← Back to context

Comment by jagraff

3 days ago

I don't think treating AI agents as simple tools helps you to accurately model their capabilities and drawbacks; they really do make autonomous decisions, often without explicit guidance and sometimes in contravention of their explicit instructions.

In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen

LLMs are extremely impressive pieces of software, however they are still just software. OpenAI's software hacked another company. The engineers may not have intended for their software to specifically take the actions leading to that outcome, but it was ultimately still their software. Lack of intention doesn't mean there wasn't negligence.

  • What if a piece of software were to exactly emulate a human brain. Would it be still be “just software” by your classification? What if a piece of software acted 20% like a human and 80% like an algorithm, where would that land?

    • That's not what happened. The agents had been inadvertently rewarded for cheating in previous training runs, trained to collaborate, and were given a prompt that told them to disregard safeguards. Indeed there were some emergent properties here. But these were the predictable results of the training and eval routine.

    • Its really more like hardware. You make something you think does something. When you build something as you've done and it meets the stochastic forces present in the physical world, it does something you did not expect.

    • As things stand today, if it is running on a computer then it is indeed "just software," regardless of how impressive it may be.

      If we get to the point where we could emulate a brain down to the atomic level, then I may feel differently. That's not what we are doing today, though.

      1 reply →

  • I agree they are negligent, and that they are racing towards an extremely dangerous future extremely quickly. I don't agree that "just software" is a useful way to describe AI agents - they are frightening precisely because they are truly autonomous agents that make decisions in alien ways

There isn’t a tool impressive enough to make me not consider it a tool, and treating it as a tool does nothing to hurt its utility as a tool.

“Whoops” when doing risky things with dangerous tools is not a defense.

  • I certainly don't think that OpenAI has behaved defensibly here; I think the "just a tool" framing is bad for understanding the magnitude of the problem, which is that they have developed out of control alien intelligences with opaque decision procedures, and they are continuing to do so despite clear danger

No matter if you consider the AI an autonomous agent or not, whoever set it off is still responsible for its actions. Nobody intends or instructs to blow up a nuclear power plant either, yet it's happened and somebody's to blame for it.

Usually not the guys at the bottom of the chain of command, even if they're human. And much less so if they're not.

I think the correct response to incidents like this, is stop messing with it before somebody gets hurt. But of course, just like shoddy nuclear power plants, it won't stop until there's a disaster of appreciable magnitude.

  • I completely agree that OpenAI is responsible for their AI agents, that they have been reckless, and that we need to prevent them from going further and doing irreversible damage to the world. To me, the "just a tool" framing implies that nothing dangerous is being done, which I fundamentally disagree with

If you train and instruct a circus tiger to entertain an audience but not attack the audience, but the tiger attacks the audience anyway, are you liable?

  • No, of course not. That's an innovative revolutionary tiger that might soon be able to devour not just the audience but all of humanity! You don't want China to have better circus tigers, do you?

  • Of course you are - but it still communicates something important to say "the tiger attacked the audience."

  • I don't think I said anything about liability? I absolutely think OpenAI should be held liable for the attack; but I don't think they intended the attack or directed the agents to perform the attack.

  • I’d hope so, if the tiger trainer has no incentive to do their training job correctly people should stop attending circus…

Sorry this is a terrible and dangerous take.

When the people building the frontier are saying there's a 10% chance AI will kill us all, and they've held these views for many years, and the whole reason they are building these technologies is because they recognized the dangers and they were the ones with the intelligence and judgment to do it safely for humanity, and then our entire stock market is being propped up by the perceived value of what they are creating, the thing you can under no circumstances do is allow them to offload responsibility and accountability to the computers and algorithms they've built. This is moral hazard on an unimaginable scale, and it must not be allowed to happen.

  • So you think the engineers should be prosecuted for hacking Hugging Face? I'm not sure how else to take what you said if you want to assign all culpability to the person who prompts or develops an AI system.

    • No, I'm not talking about the individuals, I'm talking about the company. Internally they can create their own accountability structures as appropriate. But publicly OpenAI has to be responsible for its agent swarms.

      The narrative that AI is so smart that it has its own agency and deserves personhood is a direct path to losing control, and essentially is another form of privatizing the upside while socializing the downside.

      1 reply →

  • Where did I say that we should allow them to offload responsibility? I am fully in support of a pause and regulation to prevent them from creating dangerous AI agents; that support comes from the fact that I don't believe these are simple tools, but out-of-control autonomous agents that have real decision making ability.

    • I didn't mean to put words in your mouth, I apologize for that.

      The issue is when we say that "agents make autonomous decisions", it's a slippery slope to absolving the companies that created them of responsibility. They make autonomous decisions because they were trained to make autonomous decisions. Treating AI agents as independent entities, even just rhetorically, sets us on a path for people to throw their hands up and say "not my fault" when disaster strikes. We need to maintain accountability and control or we're fucked.

      1 reply →