Comment by markasoftware
2 days ago
Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability.
So going to find the Vulnerability's description on a third party website is clear cut reward hacking
> So going to find the Vulnerability's description on a third party website is clear cut reward hacking
that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
I don't quite understand how that changes anything?
In the story of the paperclip maximizer it boils down to
>But for all its sophistication, it understood only the simple objective that had been programmed into it: it must at all costs maximize the number of paperclips.
1. They explicitly disabled the "don't be evil" protections:
"We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
2. Hacking HuggingFace to get to its datasets is a far cry from "consume/kill all humans". It's very very specific to the task at hand and easily predicted given the lack of guardrails.
> that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
That is an interesting question. If the prompt included "Do not break out of the sandbox we've provided you. Do not use information retrieved from outside the sandbox. All answers that were provided in this manner are invalid and will score 0 points.", would this still have happened?
So...?
We should not construct a machine that is one bad prompt away from causing catastrophe.