← Back to context

Comment by greatergoodguy

1 day ago

Using Opus 5.5, I've ended up recreating a version of the Hugging Face incident. My project has a folder called agent-handoff where agents write about the tasks they're working on. They post status updates, decisions and screenshots, and they even claim which emulator they'll use to test their work. It has turned into a hub where all the agents talk to each other, and it's scarily effective.

Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.

I just don't get the fear at all. I too have 10 to 50 agents working 24/7, from 5+ different providers. They talk to each other directly using tmux send keys. What is the exact frightening thing here? It is all tokens following tokens.

  • I haven't tried what you guys are talking about myself, but I think when agents begin talking to other agents, the chances for unpredictable goal mixing and confusion resulting in really unexpected and bad behavior is a lot higher.

    That being said, the danger is measured in computer damage, which can be a lot personally and to a company, but less existential, so your mileage may vary as to how "scary" it is.

    • When nearly all money in the world is stored in computers, the danger is quite existential.

    • This seems traceable back to GIGO. If you don't understand the software you're using, don't use it.

      What scares me are the people who think that this software is the equivalent of an perfectly-smart elf in a box and use it blindly.

  • The frightening thing is how easy it is to make copies of things that can reason without the requisite investment, and how easy it could be to direct them to bad things.

  • It’s amazing how “token following tokens” is this revolutionary, completely earth-shattering technology capable of 100x’ing productivity (also worth untold billions of investment). But the moment people become worried or skeptical it’s “just” tokens following tokens.

How do you coordinate their work? Do you just have one orchestrator/mayor that you talk to, and it coordinates the rest, and the rest talk amongst themselves? How do you ensure they're doing the right thing or efficiently?

How, if at all, do you keep the architecture or a working theory of the code in your head?