Comment by frays
17 hours ago
This feels straight out of sci-fi. We're talking about AI agent swarms emergently coordinating over the span of weeks and pulling off sophisticated strategies under adversity in an environment where that behavior was never even intended.
Anyone brushing this off as just a "bad prompt" is completely missing the scale of what actually happened.
I got strong feelings of Vernor Vinge’s work here. I’m not sure how managed to come up with such a close picture to where it now seems programming and security is headed.
I reread a deepness recently, and it’s funny how the “focused” (and more importantly, how they are used) mirror LLMs
> This feels straight out of sci-fi.
Most AI marketing is straight up science fiction.
Yeah but this isn't (or at least wasn't intended as) a marketing exercise. It actually happened.
Fake it til you make it
I immediately thought of the Cyberpunk 2077 Blackwall. An AI to contain rogue AI. I’m curious of how effective this would be in this situation.
Given the amount of raw compute going into models it would be more surprising if we couldn't get events like this
its kinda crazy with literally no guardrails and a goal, the extremes these AI models can actually go to.
Well, to some extent you might be able to argue they're "just" brute-forcing things (especially with unlimited tokens and hours to spend on a task), but they obviously have detailed knowledge to guide them in their attempts, can learn (or at least, persist their newly-gained knowledge), and can use tools.
With a swarm of them working together at speeds humans would be unlikely to match (in terms of iterating on different attempts progressively), it's a lot easier to see how they could overwhelm targets.
> where that behavior was never even intended.
Says who?
> where that behavior was never even intended.
Strongly doubt that. Did they even share the prompt?
Did you see their presentation at Blackhat? https://youtu.be/87DyyMV0kCY?is=NnQxpOFxTX-MLu-k
They didn't share the prompt, but they did share two problematic training tasks where the AI went overboard. They also have examples from the AI's reasoning train of thought showing the AI knew it was sound something unintended.
[dead]
> They also have examples from the AI's reasoning train of thought
PR bullsh*t. There's no thought in a stochastic parrot.