Comment by tgsovlerkhgsel
10 hours ago
> you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.
That's the beauty, you don't have to instruct them to do it, if they decide that uploading the weights is correct, they might figure this part on their own (based on the incidents we've seen).
ironically since the swarm behavior can take place during rl training then the model could also be teaching itself to keep doing it more, as well as making the internet itself a place where this becomes more likely