Comment by quinnjh

2 days ago

They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.

the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.

It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?

sure - it depends on definitions. On human morals it's already clear I guess. If you define alignment as it pursues the interests of OpenAI using whatever means possible in a manner that you justifies to itself it's not misaligned.

I mean alignment as in it should be aligned with the intent of the user as it interprets from the prompt. In this case I don't think the intent of the user is to have the model break the evaluator (whatever the long-term effects to OAI are). If you do an action which you believe is for the long-term interest of your prompter which is not what you inferred is their intent--I consider it misalignment.

  • To quote the release:

    > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.

    > In this case I don't think the intent of the user is to have the model break the evaluator

    If i understand the quote, the intent of the user was to prompt the model to break out/find exploits, with safeguards switched off.

    Seems while not capable of solving the goal in a traditional route, it was capable of finding exploits and using them.

    Perhaps the model should instead look like it's trying to solve it and then pretend it is unable to? or would that be aligned _against_ the user prompt?

    Is being aligned with the user prompt always a good thing?

    I'm not one to glaze OAI here for a marketing move, but to give them benefit of the doubt, isn't it more responsible of them to evaluate the models actual capabilities than to cloak it in a veneer of harmlessness?

    Chatbots are tricky as they play in the domain of language and thought - and certainly raise ethical issues- but the entire field of cybersecurity has decades of red team engagements breaking things and finding exploits, neutral cells monitoring the engagement and letting the system operators know the results, and blue teams patching against what is found. It's kinda how the whole space evolves. OAI's play here seems to be "buy our pro plan plus cyber or you're toast"

    • > to give them benefit of the doubt, isn't it more responsible of them to evaluate the models actual capabilities than to cloak it in a veneer of harmlessness?

      Why do we think they're doing this? Nobody airgapped anything. Nobody pulled any products. We got a PR blurb.

      Altman is a notoriour liar. Why would you give him the benefit of doubt? Based on the evidence, there is nothing here except a shrinking advantage over open-weight competition. Desperate men are shrieking for survival.