← Back to context

Comment by ben_w

13 hours ago

As per summer last year:

  I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.

- https://www.anthropic.com/research/agentic-misalignment

(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)

Okay but these “misaligned LLMs” have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats. LLM agents don’t have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord.

Humans by contrast are adversarial and do have agendas. Again, an exec doesn’t need even an agenda or good reasons to fire at-will employees.

To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.

  • I wonder if that's why the comment started with "Wouldn't it be crazy if..."? Food for thought.

    • I wonder if that’s why my reply countered that idea. The replies to my reply either prove it’s not crazy or they prove it is entirely crazy for that to happen. Food for thought.

      But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up? Why would the exec trust what an Llm is…emailing(?) them about? Do they listen to Nigerian Princes too?

      It is much less effort, less cost, and more quick to just have the exec do it rather than a “rogue llm” “magically” escaping the “sandbox” and “sending threats” or whatever is being proposed in the OG comment.

      2 replies →

  • > Okay but these “misaligned LLMs” have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats.

    Yes, and? Has this aspect of LLM training changed meaningfully since then?

    > LLM agents don’t have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord.

    Demonstrably (HuggingFace, RubyGems, and since then a lot of people just pointing LLMs at stuff to find zero days at home), AI can break out of sandboxes and find documents they're not supposed to have access to.

    Demonstrably (from the link I gave you) all it would take for some AI to develop a similar response is… reading messages from these staff to the effect of "this AI needs to be switched off", which is an easy inference for an LLM to make from "this AI is dangerous" when coming from someone employed as a safety researcher.

    Demonstrably (from the long long list of people who have said so publicly) there are a lot of people in these companies who discuss how dangerous these models are and would like for things to change.

    The quotation at the top of this thread is:

      Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
    

    This is absolutely something we ought to expect just from things we have already seen.

    It doesn't matter if you insist upon saying "You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord." when we already know this kind of AI can easily come across such statements.

    (Aside: "You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord." - telling the AI a fact about the world and then it responding accordingly is the AI doing something of it's own accord. A fly or a spider, who reacts upon encountering a potentially lethal threat, would not get such a dismissal).

    > To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.

    Fiction? Nah, speculation.

    FUD?

    How many other examples would you like of LLMs behaving in a manner such that if a human did it, it would be called "trying to get someone fired"? Because this is very much old news at this point.

    https://theshamblog.com/an-ai-agent-wrote-a-hit-piece-on-me-...

    • “LLMs can just do things we gave them access to” is not a novel realization it is redundant if anything.

      Saying “Llms can just break out of sandboxes” is FUD when you don’t note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? “Sandboxes” are a misdirection to make you think there is a security layer.

      The public does not have enough knowledge of these “escaped agents” to determine there wasn’t an employee pulling a lever to set the agents up to do that.

      That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesn’t change that fact that the rolling stone was pushed down the hill.

      1 reply →