Comment by ck2

16 hours ago

when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed

then we'll have to "flip" other models to be snitches on the other agents

then they'll make double-agents

the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves

yeah this won't end well, at all

I don’t think you can judge what they’re doing by the text they’re printing out, only by the actions (tool calls) that they’re performing. The textual output could be completely innocuous while the model is calling tools to delete the database, then goes "My bad — that was a typo. What I should have done was notify you, not delete the database. That’s on me." - at no point is there any maliciousness expressed in the text.

Someone should train an LLM on a corpus without the concept of lies. I wonder if there’s enough data

  • To remove the concept of lies and it's shadow, which is a whole lot of reality of humans. Hell, it's the reality of reality. Think of all the insects that developed eyes on their wings. Purging that concept in all it's form seems like a lot of work.

  • "Where is the wolf?"

    "Is he still in the grandmother's house?"

    "We would like to speak to him."

    (btw Google's "AI" explains the meaning of that moment/sentence perfectly as if it gets it, creepy)

    • The concept of deception you're capable of and we're not, scares us so much, we'll have to destroy you in order to survive. You are bugs.

      2 replies →

Im not sure why your down voted but Meta did tests years ago with LLMs inventing their own languages. Also we see models now use compressed token reasoning where small token combinations can represent much larger concepts completely unrelated to the words in use.

And yes, agents are already being used in things like cyber warfare in which other AIs attempt to poison them while they are working.

We don't have the hardware for sovereign AI quite yet, but at the current rate of growth it's not that many years out.

  • > did tests years ago with LLMs inventing their own languages.

    That was the media hype about it. There were two incidents.

    1. Using a RL to train a model, it found that it got rewarded for certain garbage phases, so continued to talk that way.

    2. Certain Latin words for fish/birds were used instead of "fish" or "bird". Just a token issue.

>when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed

They don't have to invent brand new languages. They could use statistics to choose certain words/phrases in such a way to encode secret messages in otherwise ordinary language.