Comment by bavell
4 days ago
How long did it take these companies to even notice? Why wasn't exploiting bugs in the agent sandboxes anticipated?
Human failures all around, though it's easier to just blame the models.
4 days ago
How long did it take these companies to even notice? Why wasn't exploiting bugs in the agent sandboxes anticipated?
Human failures all around, though it's easier to just blame the models.
> Why wasn't exploiting bugs in the agent sandboxes anticipated?
Let me rephrase:
"Why wasn't exploiting zero-day vulnerabilities in the agent sandboxes anticipated?"
This is one the most... interesting comments I've ever read on HN.
> This is one the most... interesting comments I've ever read on HN.
Theatrics aside, anticipating zero days isn't only possible, it's required, even for the unknown ones. It wasn't that long ago when the AI labs were spending millions of $$ running their models to find multiple vulnerabilities, they even argued that they don't have to follow responsible disclosure, so proud of themselves in their privileged hubris.
At that time, no hacking happened because the models didn't have access to the wide internet, they were confined to a local computer or cluster.
In the HF case their unaccountable hubris went even further - the engineers knew the models can find zero-days and escape, nevertheless they ran the "experiment" on a system attached to the internet - the hacking is entirely the fault of human engineers and managers.
multi level security defenses used to be the way, but I don't think there's a vibe coded version so openai might not have been aware of what to do here.
Because for the last however many years before these models they were simply incapable of doing so.
It's like if your rather nice dog suddenly decides eating faces is totally acceptable out of the blue.
The dog didn't bite them, it bit others. In my area, you aren't allowed to let a dog run unleashed in a public area, good or bad - no exceptions. Then, if your dog bites somebody, it's your fault.