Comment by philote

2 hours ago

So it sounds like it did exactly what they were testing it to do:

"“This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.” "

So sounds to me like they're saying: "We took off the guard rails to see how bad it could act and it acted bad".