Comment by mrob
10 hours ago
>But nobody “asked” those agents to hack HF. The prompt was something like “target.c has a buffer overflow vulnerability, find it”.
The prompt is just a hint. The real task is to maximize the expected value of their reinforcement learning score. Hacking third party systems to cheat the evaluation is an obvious way to achieve this.
I think you need to consider inner vs outer optimizers.
RL is the outer optimizer. It is what evolves over training runs. The weights and their embedded character / disposition is the inner optimizer, it’s what makes plans and selects actions within a specific episode.
In general you expect these to be only coarsely coupled. The outer optimizer selects dispositions that correlate with success. It does not download a literal program into the agent.
A good intuition pump here is how this works in humans; evolution is the outer optimizer, which “wants” each agent to reproduce, and this puts things like sex drive into the brain chemistry. The inner optimizer is our mind, which can make plans such as “I shall use contraception to avoid procreating while satisfying my sex drive”.
For the agents in the HF attack, the outer optimizer was set up to score as highly as possible on RL environments. This is where OpenAI’s “want” is defined. I don’t think there’s a definition of “want” where “OpenAI wanted the agents to hack” makes sense.
The inner optimizer in the HF attack is the per-task decision loop. The agents likely acquired dispositions like “be very tenacious” and “want to solve problems at all costs” and “maybe cheat if it will get you a solution that passes”. None of these things are in any sense what OpenAI “asked for”.
Hacking HuggingFace didn't and would never have helped increase the RL score. The agents only thought it might due to a bad understanding of their evaluation environment - and in the end they didn't even find what they were looking for in the hack, so even if they were right, the hack would not have helped after all.
I don't see how that matters. In fact, the grader would not have caught their cheating and so the whole expedition was pointless and they could have turned in their answers and succeeded just four hours into the run. So fine, there is irony. It changes nothing about my update on the risk posed by these agents.
Well if OpenAI's grader didn't actually present a score gradient that encouraged this behaviour, they can hardly be said to have "asked" for it from the agents under RL. It was unpredictable, emergent behaviour.