← Back to context

Comment by amelius

2 years ago

Guessing. I think they may have trained on copyrighted stories. So what they supposedly did was ask an llm to replace the names of the main characters. Since Elara is not a frequently used name, there is little chance of clashing. Then they trained chatgpt on those stories.

Has any legal precedence been established that training breaks copyright? Are you implying that reprinting a novel word-for-word, except for changing one character's name, wouldn't be copyright infringement?

Highly unlikely. The more likely truth is just that the default temperature of ChatGPT is really low (but not quite zero); so it keeps spitting out (roughly) the same story. It does switch up protagonists a bit and different versions of ChatGPT also have different names for them. Slightly modified prompts also return different names (eg. "Tell me a sad story").

when asked the reason, ChatGPT had this to say- "Actually, the choice of “Elara” wasn’t a result of training on specific copyrighted stories or any prompt to avoid copyright claims. OpenAI models like me are designed to create original content without directly referencing copyrighted characters, and "Elara" is simply a popular-sounding name in many storytelling contexts. I just used it consistently for its versatility, but I’m totally open to switching things up!"

  • ChatGPT's opinion on the matter is completely worthless, unless it was also trained on an accurate description of its training process (it wasn't). Language models do not even have access to their own "thought process" - if you ask it "why" it said something, you will get a post-hoc rationalization 100 percent of the time because the next-word prediction only has access to the same text that you see. The rationalization might be incidentally correct, or it might not - either way it contributes no real information about the model's internal state.