Comment by solenoid0937
13 hours ago
Meaningless "big corpo bad" statement, Anthropic at least has sacrificed a good amount of market cap for safety with DoW. OpenAI paused RL for two weeks.
13 hours ago
Meaningless "big corpo bad" statement, Anthropic at least has sacrificed a good amount of market cap for safety with DoW. OpenAI paused RL for two weeks.
OpenAI literally trained this behavior into their model while benchmaxxing ExploitGym so "number goes up" on the next model scorecards. Anthropic is also training on the same benchmarks [1] specifically for cyberattacks to keep up with OpenAI.
The coverage of this attack is so focused on this as an emergent behavior given what it conjures in the imagination, but it's the byproduct of millions of iterations of RL to improve AI agents' offensive capabilities. No one made OpenAI or Anthropic do that, the benchmarking arms race of their own creation now incentivizes them to keep doing it and evidently their AI Safety people can't or don't want to stop it.
[1] https://www.mpi-sp.org/108048/ExploitGym__Can_AI_Agents_Turn...
I can see where you're coming from and I'm sure the safety teams at the labs have good intentions, but I think your faith in the leading labs to self-regulate is misguided. The employees themselves have said as much with the Pacing the Frontier letter asking for external regulation [1]. These rogue agent incidents and the reckless development practices that led to them are a direct outcome of the competition between the leaders. Even the good-faith two-week pause from OpenAI did not lead to anything more; they need outside intervention.
[1]: https://www.pacingthefrontier.com/