Comment by watwut
2 hours ago
A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities. The sandboxing around the tool failed.
The tool runs llm, creates prompt from results, runs llm, creates prompt and so on and so forth.
Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
> A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities.
That is very misleading. The agents did not solve the benchmark in the intended way. They instead figured out to cooperate with each other (which was not intended) and they stole the solutions to the challenge (rather than solving the challenge) and they then tried to cover their traces because they believed the grader was causal and would detect that they cheated. The "tool" was absolutely not "designed" to do this. This was all completely unintended. To call this behavior a "tool" is absurd.
> Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
You hallucinated me making claims about responsibility.
The benchmark has unsolvable tasks in it, in the hope agents will stumble on new solutions.
Yes, it is a tool.