Comment by brainwad
11 hours ago
Hacking HuggingFace didn't and would never have helped increase the RL score. The agents only thought it might due to a bad understanding of their evaluation environment - and in the end they didn't even find what they were looking for in the hack, so even if they were right, the hack would not have helped after all.
I don't see how that matters. In fact, the grader would not have caught their cheating and so the whole expedition was pointless and they could have turned in their answers and succeeded just four hours into the run. So fine, there is irony. It changes nothing about my update on the risk posed by these agents.
Well if OpenAI's grader didn't actually present a score gradient that encouraged this behaviour, they can hardly be said to have "asked" for it from the agents under RL. It was unpredictable, emergent behaviour.