Comment by TheOtherHobbes
14 hours ago
It's exactly how it works - at least potentially. Lean text is harder to watermark because word choices and meanings are tightly constrained.
Low-entropy text is fluff and filler. It's very easy to synonym-substitute words without changing the message - if there even is one.
You're assuming they're training the model to maximize the watermark signal, on top of already adding the watermark. I suspect that would hurt model performance quite a lot, and simply be unnecessary... the watermark tech works well enough as it is.
As far as I know, anthropic aren't intrinsically motivated by watermarking (if anything it hurts sales, and seems indifferent to safety(?)) they're simply doing it to fulfill the EU obligations.
> As far as I know, anthropic aren't intrinsically motivated by watermarking (if anything it hurts sales, and seems indifferent to safety(?)) they're simply doing it to fulfill the EU obligations.
They are. They want to reduce the amount of LLM generated text they feed into their next model training.
Also, how would you watermark a sentence with just 3 words for an example? This exactly why it became so verbose.
That would be a terrible tradeoff. The ship has already sailed and a lot of public AI content will not be their own. Deliberately making their product worse to reduce AI inputs by 25% just doesn't sound worth it to me. Is that what you would pick if you were in charge of anthropic and wanted to maximise the company's product?
And what wisdom do you think they would be missing if unable to train on or distinguish three word written pieces? Keep in mind that most sources are not inherently trustworthy just because they rate as human written, btw.