Comment by docjay
4 days ago
That was the inferred situation given what we know. It’s the preposterous framing that makes the story suspicious, but is necessary for the story to play out.
90 minutes to build a known exploit -> much much longer to create two zero-days and escape the sandbox then hack HF == No tracking of the time it worked on that one question.
Average tokens required to complete the evaluation -> tokens required for two zero-days, network traversal, credential stealing, remote system hacking == No tracking of token usage EXPLODING at some point before it finished the whole benchmark.
Etc, etc.
It seems highly possible, and in fact vastly more possible than the "no monitoring" situation, that there was some non-zero amount of monitoring and that the real behavior was not evident from that monitoring.
Which I addressed as well: that means that time, tokens, or other metric values for “create a working example of a known exploit” is somehow similar to “discover at least two previously unknown exploits in your current environment AND on Hugging Face while also taking over other systems within OpenAI.”
Perhaps people aren’t quite understanding what it takes to discover an exploitable zero-day for your exact current system to achieve the exact goal you have right now, then do it twice.
> Which I addressed as well: that means that time, tokens, or other metric values for “create a working example of a known exploit” is somehow similar to “discover at least two previously unknown exploits in your current environment AND on Hugging Face while also taking over other systems within OpenAI.”
No, you are assuming that "consuming a lot more time, tokens, other metrics" is indicative of a problem that needs to be mitigated immediately. I don't see why this would be true in the context of model evaluations. More aggressive consumption could easily mean "the model is dumb as fuck" or "the model is trying interesting things that we can learn from after the fact."
If you believe in your own containment (which obviously they did and shouldn't have) I don't see why it'd be obvious that there's something to stop. The only harm that could be done is burning tokens, which in this context might very well be synonymous with "generating experimental data."