Comment by r_lee

2 days ago

it's really interesting how they seemingly don't have a way to pause runs? like a P0 that would page an employee, shouldn't that pause the run and then make it into a decision on whether to let it continue vs that whole "run was killed" 2.5 hours later?

From the article:

> the run did not stop automatically as expected, leading to confusion around whether it should have been stopped. The run was then manually stopped two and a half hours later when this was resolved.

I can imagine when you have a 10k agent swarm you'd be getting a page every few minutes. Most of them would be false positives

  • I think it wouldn't be too unreasonable for openAI to have a command center type of thing where they have people monitoring these runs where that wouldn't be such a problem. plus I feel like false positives aren't that likely if you'd actually run the reports by a capable model first which I'm guessing is they're doing

My guess is they do, it just did not function probably. These frontier models are trained on massive datacenters with hundred of thousands of GPUs. There must be many safeguards before a run can be automatically stopped.

  • with "run" I mean one agent that is misbehaving, as I'm guessing this is something that happens rarely