It literally re-weights the output tokens from what the LLM would otherwise have chosen. It _has_ to. It can't be positive, because then that's not watermarking, it's a better LLM.
What if the next token represents a wrong or low-quality answer, but would have only been picked 10% of the time, but now it's picked 20% of the time? Doesn't that obviously decrease the model quality, even though "it might have picked that token anyway"?
Unless you’re at 0 temperature, there is no single token it would have chosen. It’s always picking one of multiple randomly according to a probability distribution.
It literally re-weights the output tokens from what the LLM would otherwise have chosen. It _has_ to. It can't be positive, because then that's not watermarking, it's a better LLM.
It's a very unintuitive algorithm, and is pretty clever.
I recommend reading up on it: https://www.nature.com/articles/s41586-024-08025-4
But no, it only ever picks tokens that are in the probability distribution of the last layer, and it might have picked anyway.
What if the next token represents a wrong or low-quality answer, but would have only been picked 10% of the time, but now it's picked 20% of the time? Doesn't that obviously decrease the model quality, even though "it might have picked that token anyway"?
Unless you’re at 0 temperature, there is no single token it would have chosen. It’s always picking one of multiple randomly according to a probability distribution.
Give me an example how would you watermark a single short sentence like "I like turtles"?
1 reply →
Unless you’re running at temperature 0, there’s not one single token that the model definitely would have chosen each time.