← Back to context

Comment by derangedHorse

4 days ago

> The way OpenAI seems to want

This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light.

> The second seems to forget that jailbreaks are available for every model

Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks.

This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to.

You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true:

"2. OpenAI’s harness and network security controls were unintentionally [...] bad"

The post was long enough so I couldn’t capture all the nuance and details for sure. Also, this comment was an opinion based on limited info right now, that may change if we found out more. I think OAI does want it framed this way but that’s something we’ll likely never prove if it’s true.

Your comment about jailbreaks being more one off and hard to do consistently in agents is a good point. Still getting an agent to hack isn’t hard even without a jailbreak, you just have to tell get creative in what you tell it. I’ve found telling it that it’s in a CTF or that I own the system that it’s hacking will work fine. A lot of offensive security companies are running agents in their testing so getting an agent to hack seems commonplace.