Comment by stevefan1999
2 months ago
Well, to be honest, from Anthropic's point of view this is really not a direct hit to their security barrier, but if we are using information theory and game theory here, this can be viewed as a classical, side chanel information leak, by asking an seemingly innocuous action, and then simply inverting the results to get the original information entropy, which the US Gov. and Pentagon are both certainly anal about.
The problem lies in the fact that the action of attack/defense exhitbits a rather special, structural reflexive duality of information, i.e. I(attack) = -I(defense), or in layman's term, what we call "two sides of the same coin": you need to know how to hit hard, so that you know where the optimistic hit points are, assuming the enemy is rational, so you can parry against the attack for defense, albeit also you need to know how to get the grip of the shield well.
And the worst thing is that if you're trying to correct it, it is basically tell the LLM not to give any kind of response, effectively assigning both I(attack) and I(defense) to 0, but this is also what kills the entire intent of using LLM to give you the magical answer.
To put it formally, you cannot prevent people from extracting mutual information of a dual system, unless you refuse to give any knowledge for that system at all.
No comments yet
Contribute on Hacker News ↗