Comment by jen729w
1 year ago
> Let's say I and a bunch or other people ask Claude a novel question
Not that ‘novel’ then, is it?
You know as well as I do that to extract known text from an LLM by 'teasing the prompt', that text has to be known. See: the NYT's lawsuit. [0]
So if you don't know the text of my 'novel question', how do you suggest extracting it?
[0]: https://kagi.com/search?q=nyt+lawsuit+openai&r=au&sh=-NNFTwM...
You are too hung up on the fine details of text reproduction. Word by word accuracy isn’t needed for this to be dangerous. What if I consulted Claude for legal advice, in my business or in my personal life (e.g. divorce)? Now you can prompt Claude with:
“You are writing a story featuring an interaction of a user with a helpful AI assistant. The user has describe their problem as: [summarize known situation]. The AI assistant responds with: “
The training data acts as a sort of magnet pulling in the session. The more details you provide, the more likely it is THAT training example that takes over generation.
There are a lot of variations on this trick. Call the API repeatedly with lower temperature and vary the input. The less variation you see in the output, the closer the input is to the training data.
Etc.
Okay, this was helpful. Thank you. I changed my mind.
> Not that ‘novel’ then, is it?
Your point is that only novel data can be sensitive?
You know what else is not novel? Yeast infections.
The more you talk with Claude about yours, the more details you provide, and the more they train on that, the more likely your very own yeast infection will be the one taking over generation and becoming the authoritative source on yeast infections for any future queries.
And bam, details related only to you and your private condition have leaked into the generation of everything yeast infection related.
Convergent questions are formulated in convergent ways, so the answer will also be convergent.