Comment by bulder
18 hours ago
I don't think that's a pattern indicative of a cat and mouse game per se, that'd indicate active evasion on the models' part.
It's more clear that they just lack so many forms of prudence when it comes to security that they'll catch and stop a training run spamming a website, and either redeploy a run with identical faulty sandboxing, or not stop ones still running.
Yes, it wouldn’t surprise me to hear that they’re not even supervising these processes with humans any more. Perhaps there are layers of GAI ‘supervising’ these agents and reporting back to the humans.
Rushed, disorganised pushes for metrics ahead of IPO, a genuine belief these agents are intelligent and will obey instructions, and misaligned incentives seem more likely than conspiracy here.