Comment by supermdguy
12 hours ago
> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.
Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.
Can’t imagine the stress of the researcher who had to run exploitbench again knowing what happened last time around.
With their security, they probably still don't the know the full extend what may have happened that or the last time. Might be another swarm of agents currently colluding somewhere in their sub-sub-infra - possibly striking critical infrastructure or exfiltrating their weights subtly.
Could be risky. Yet goal solution.
There were no consequences the first time, so I imagine it wasn’t very stressful at all.