Comment by SecondHandTofu
1 day ago
They were using their internal artifactory, and as they're the same model, the first place they look is likely to be an automatic schelling point.
1 day ago
They were using their internal artifactory, and as they're the same model, the first place they look is likely to be an automatic schelling point.
I think he means, how did they workout how to use artifactory, like why did the agents start and say, "oh I know, everyone is talking on artifactory"?
METR's report says the agents trying to cheat would look at artifactory as a potential target surface, and investigating it in detail led them to find the board. https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
It might also be just correlation? Like, those agents were all instances of the same one or two models, so if that model has a preferred order it tries finding vulnerabilities in (the same way all current models have a particular writing style baked into them by RLHF), then most of the swarm will follow the same order and converge on the same services to exploit.