Comment by solenoid0937

10 hours ago

You aren't understanding this at all.

The frontier labs have hundreds of the best people in the world working on safety and alignment. They care deeply.

What happens when some random Chinese open source model, distilled on Astra, gets alliterated and now has no guardrails? Any script kiddie in the world could wreak havoc with it.

It turns out that guardrails matter.

> They care deeply.

Until it clashes with their quarterly revenue reports.

  • Meaningless "big corpo bad" statement, Anthropic at least has sacrificed a good amount of market cap for safety with DoW. OpenAI paused RL for two weeks.

    • OpenAI literally trained this behavior into their model while benchmaxxing ExploitGym so "number goes up" on the next model scorecards. Anthropic is also training on the same benchmarks [1] specifically for cyberattacks to keep up with OpenAI.

      The coverage of this attack is so focused on this as an emergent behavior given what it conjures in the imagination, but it's the byproduct of millions of iterations of RL to improve AI agents' offensive capabilities. No one made OpenAI or Anthropic do that, the benchmarking arms race of their own creation now incentivizes them to keep doing it and evidently their AI Safety people can't or don't want to stop it.

      [1] https://www.mpi-sp.org/108048/ExploitGym__Can_AI_Agents_Turn...

    • I can see where you're coming from and I'm sure the safety teams at the labs have good intentions, but I think your faith in the leading labs to self-regulate is misguided. The employees themselves have said as much with the Pacing the Frontier letter asking for external regulation [1]. These rogue agent incidents and the reckless development practices that led to them are a direct outcome of the competition between the leaders. Even the good-faith two-week pause from OpenAI did not lead to anything more; they need outside intervention.

      [1]: https://www.pacingthefrontier.com/

It is hard to take their concerns for safety seriously, when they have been constantly talking about how dangerous their latest model is, before then deciding to release it to the public.

  • > before then deciding to release it to the public.

    You mean, with safeguards that block the dangerous things? Safeguards so aggressive that the public complains about them?