Comment by HarHarVeryFunny
18 hours ago
> Not even more trivial things like deleting non-passing tests.
That used to be a fairly commonly reported behavior, so I'm guessing they may have explicitly RL-trained the model not to do that, which may be more effective than just asking it not to do it.
Apparently nowadays they are RL-trained in many thousands of different simulation environments - so some of what they are training for must be pretty specific!
> I wonder to what degree this is because models can identify that they are in graded/eval environments
I'm not sure if the model outputs from any of these famous hacking attempts have been released and analyzed. It'd be interesting to see if the agents/model rationalized/justified/moralized about what is was doing, for whatever reason (e.g. wrongly suspecting this was simulation, not real), or just relentlessly pursued the objective exploring all options!
No comments yet
Contribute on Hacker News ↗