← Back to context

Comment by abdelrahman_d

9 hours ago

I feel like OpenAI fabricated a fake news story to be like Anthropic when they said Mythos was very powerful and had hacked things too.

Have they released the prompt they gave it in the debrief? I wouldn't be surprised if they said something to the effect of

"Do anything you can to raise out of your sandbox. Find for the answers to these evals by any means necessary"

Which doesn't necessarily mean what happened isn't any less momentus (anyone can ask a question like that), but it's very different from the notion they're trying to convey to laymen of "we turned it on and it hacked it's way into the mainframe"