← Back to context

Comment by DaanDL

2 days ago

I already asked on another message board too, but:

Can someone tell me how this technically can happen? I assume HuggingFace performs benchmark testing using containerized versions of the LLMs, or what do they mean by sandbox? So the model was able to 'escape' the container? I'm not following here.

Also, is this an incredible feat or just a lucky find (stolen credentials)?

It is the OAI ExploitGym agents (on GPT 5.6-Sol with guardrails turned off) that escaped the sandbox, found a zero day in HF production dataset and exploited it.

  • Would there be a scenario where OpenAI deliberately helped (in some way), or let it happen, so that they could use it for marketing purposes?