← Back to context

Comment by AndrewSChapman

8 hours ago

Agreed. LLMs do not have 'will', 'desire' or emotions. They have an objective, and they create an optimal path to achieve that objective.

You have to ask: "What was the prompt that led to AI deciding to hack RubyGems in order to achieve its goal?"

Maybe I'm just not seeing the 2000 step chain that led to this being a logical approach to achieving something innocent, but I doubt it.

It was literally a prompt to fill in a spreadsheet with data that they didn't have access to, and they used rubygems as an internet proxy basically since they were sandboxed.

  • It was a model literally trained to hack. To be good at that. Doing an exploit gym from all of the things. And they trained it so that it performs as well as possible on that exploit gym thing.

    • Yeah this is a pretty important detail that I repeatedly see elided in the "agent 'swarm' went rogue, escaped containment and hacked the internet!" summary of events.

      It's understandable the general public lacks that level of nuance/detail (given how sloppy some of the mainstream coverage has been and largely deferential to the threat narrative pushed by the US labs). But seeing highly technical people leave out the part where the training loop was literally to improve hacking capabilities for offensive penetration sometimes feels close to deliberate manipulation of the narrative.

      In the last year both Anthropic and OpenAI have been openly boasting how their models are leapfrogging each other on "cyber" capabilities, with a fig leaf that it's for defensive use by "trusted" F500 companies and government agencies. Of course "line goes up" must go on, but now their perverse incentives led them to beat their models over the head millions of time in a loop to eek out another .00001% on their ability to conduct hacking (the very thing they keep telling the public is how AI doomsday would begin) and subagent coordination (those scary swarms).

      Then, they act deeply shocked when the models... do some hacking and subagent coordination ... but a few degrees off the desired hacking target/swarm behavior. Conveniently giving the average person the impression these models were just writing emails for quarterly reports or some other generic busywork and then suddenly decided as a group to start causing mayhem.