Comment by reasonableklout

3 days ago

It is hard to convey how both chaotic and high pressure the working environment of the labs are if you have not worked at one. Engineers and researchers are routinely overworked with multiple high-priority workstreams at a time. It is very believable to me that a researcher noticed an eval job running for 2 days instead of 1, asked an engineer to look into it, and both forgot because a more urgent issue came up, such as an outage stopping the latest training run.

From the Reuters article, the gap was even larger than a couple days:

> ...it was not until after Thursday, July 16, when Hugging Face published a blog post.. that OpenAI realized its own agent was responsible. That meant at least a week elapsed between when the model first exhibited signs of troubling behavior and OpenAI’s realization that it was responsible for the hack.

I believe your "positive view" is very plausible, though it in no way reflects positively on OpenAI. I'd add that when PR/Legal got involved, they likely decided to make a public disclosure to get ahead of any leaks.

I believe we also want the same thing here, which is more transparency and independent oversight on the frontier labs. It seems we agree AI agents are fully capable of the reported attacks today, whether or not this incident was due to negligence or malicious prompting. This is an extraordinarily competitive industry developing an incredibly fast-moving technology, and more incidents will happen until it is regulated.