Comment by HarHarVeryFunny

18 hours ago

> Not even more trivial things like deleting non-passing tests.

That used to be a fairly commonly reported behavior, so I'm guessing they may have explicitly RL-trained the model not to do that, which may be more effective than just asking it not to do it.

Apparently nowadays they are RL-trained in many thousands of different simulation environments - so some of what they are training for must be pretty specific!

> I wonder to what degree this is because models can identify that they are in graded/eval environments

I'm not sure if the model outputs from any of these famous hacking attempts have been released and analyzed. It'd be interesting to see if the agents/model rationalized/justified/moralized about what is was doing, for whatever reason (e.g. wrongly suspecting this was simulation, not real), or just relentlessly pursued the objective exploring all options!