Comment by apsec112

3 days ago

(this is new, not any of the previous incidents)

"Investigation and response

Incident timeline:

9:50:23 a.m. The agent made the DNS tool call that received an external response.

10:02:11 a.m. The monitoring system raised a P0 alert.

10:05:06 a.m. A human reviewer acknowledged the alert.

12:34:30 p.m. The run was killed.

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions. We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior."

20 years ago I would download testing in an isolated network with no route to the internet. There was a single bastion that had two physical NICs and no route between them. It was inconvenient, but we could run massive tests that flooded switches and there was no risk to any of our other networks.

This was a precaution taken 20 years ago when we were working with stone knives and bear skins, as our load test operator would say.

If the model shouldn’t have access to the internet, put it on a network without access to the internet.

  • Another thought: openAI literally has an (almost) entire copy of the public internet they use for their training dataset! Why cannot they create an internal version of it that doesn’t require accessing public servers? They have the data already

Am I reading this right? 3 minutes to acknowledge the alert, more than two hours to act on it?

> When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions.

But they apparently won't do anything to address the possibility that simply trying to make the model behave might not work. They won't actually make sure "that the model could not access the live internet" by, for example, creating a physical hardware environment that lacks this capability.

it's really interesting how they seemingly don't have a way to pause runs? like a P0 that would page an employee, shouldn't that pause the run and then make it into a decision on whether to let it continue vs that whole "run was killed" 2.5 hours later?

  • From the article:

    > the run did not stop automatically as expected, leading to confusion around whether it should have been stopped. The run was then manually stopped two and a half hours later when this was resolved.

  • I can imagine when you have a 10k agent swarm you'd be getting a page every few minutes. Most of them would be false positives

    • I think it wouldn't be too unreasonable for openAI to have a command center type of thing where they have people monitoring these runs where that wouldn't be such a problem. plus I feel like false positives aren't that likely if you'd actually run the reports by a capable model first which I'm guessing is they're doing

  • My guess is they do, it just did not function probably. These frontier models are trained on massive datacenters with hundred of thousands of GPUs. There must be many safeguards before a run can be automatically stopped.

    • with "run" I mean one agent that is misbehaving, as I'm guessing this is something that happens rarely

A full hugging face timeline would surely make them look terrible.