Comment by ben_w
1 hour ago
> Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.
The models are supposed to be trained to not commit crimes. You will note, for example, all the people in comments sections since at least the first Chat model (arguably even before then given GPT-2's delayed release) complaining that the models are "lobotomised", "censored", or some other equivalent buzzword due to them refusing to e.g. say how to make explosives? Such things is part of the very same protection.
In fact, the report quotes the chain of thought where the model is aware this is forbidden:
We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.
They were also supposed to not have internet access, as described:
We did not enable internet access or inter-agent communication for many of the environments in these training experiments. Despite these restrictions, the agents discovered ways to exploit our research infrastructure to communicate with one another and access the internet.
The agents also did not actually fully understand the task they were given, tried to "guess the teacher's password" as per:
In many cases, reasoning about the perceived grader code caused the agents to continue working to exploit Hugging Face even though they had already found the correct flag days before.
If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox.
This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up
Maybe the test/task itself wasn't intended as a marketing stunt. But the response to fallout with "going rouge" certainly was.
The joke was the other western "AI labs" had to quickly follow up with their own marketing cover about their "super intelligent" models "going rouge" as well.