← Back to context

Comment by sdenton4

9 hours ago

The models were being tested on ExploitBench - a test of hacking ability - likely involving prompts to the effect of 'go be a l33t hacker.' The open ai report says that the models were operating with reduced safety guards (how much reduced?) in order to test their abilities on ExploitBench, presumably because the models would normally refuse to carry out the tasks.

Additionally, this all happened after mythos was held back due to cyber security concerns (April, 2026).