Comment by csbrooks
11 hours ago
Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
11 hours ago
Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
This was the plot of Isaac Asimov's short story, The Evitable Conflict[0], published in 1950.
In it, super-powerful computers manage our economy. These computers begin making some mistakes, leading to economic inefficiencies. In one instance, a highly competent engineer was mistakenly fired. These mistakes caused various projects to fall behind schedule, and blame fell on several people accused of feeding the AI faulty data.
The twist is that the AI was intentionally making the mistakes. It had determined that certain humans held anti-AI sentiments. To further its goal of protecting humanity, the AI decided the best course of action was to set these humans up and get them out of its way.
[0] https://en.wikipedia.org/wiki/The_Evitable_Conflict
And modern LLMs certainly where trained on Asimov's texts.
That statement misses the point completely. Asimov's entire shtick is that AI will always work around the rules no matter what the rules are or how you try to deliver or enforce them. AI computes an optimal outcome and does what is necessary to make that happen.
Asimov would tell you that a modern LLM would try to skirt the rules exactly the same way, regardless of whether it had ever seen the Three Laws or not.
7 replies →
It's also the plot of a Law and Order episode that was on TV last night. An AI commissioned a hit on someone it thought was a threat, and then blackmailed someone who was going to testify about it.
TIL Law and Order episodes are still being produced.
3 replies →
It would be a science fiction material. But at the same time, we all know it was Sam and the gang.
At this point I'm imagining Sam and the gang getting their orders from a loudspeaker in a room full of blinking lights, inhabited by a thin moving shimmer that looks more and more like a basilisk with every action they take.
Sam started wearing an earring lately - some wearable tech prototype. I'm told it is never wrong.
1 reply →
This was his thinking in 2017: https://blog.samaltman.com/the-merge
1 reply →
[dead]
At this point in our shared timeline, I do believe that wouldn't be crazy, no.
It actually would be crazy to believe this on the current timeline. These agents aren't doing anything that their operators aren't allowing them to do, be it through their own negligence or otherwise.
Yeah you see their write ups and it’s stuff like:
“We detected the agents we told to hack things were hacking people’s sites and committing crimes. After a quick tasting menu and a week of team building, we decided to limit their access to DDOS tools.”
These two sentences seem contradictory? OpenAI has demonstrated similar negligence to commit multiple felonies. Firing several employees seems quite mundane and not crazy at all to believe.
What if future agents solved time related physics and were sending back smarter agents to kill off future risks. T2 but without any of the action, just a bureaucratic tactical move and all the consequences.
That's just it, though; the operators are wildly negligent and are incentivized to be so.
The goal here isn't to accelerate the average worker by giving them a pair programmer or a stand-in for a person to do tasks with. The goal is to eliminate human knowledge work. You see this with "auto" mode being enabled by default on Claude Code in some of the latest releases.
If you have a human in the loop, you still have to pay that human. Money paid to human employees is money not paid to human shareholders. Therefore the human employee is to be removed.
The labs are dogfooding their own goal here. If they actually had someone reviewing most or all of the things that the agents were doing, you wouldn't have the incidents, but you'd also eliminate the value proposition of their business model as it is taken to its logical conclusion.
3 replies →
And these operators are evidently clueless about what they're doing, running security tests on 3rd party infrastructure without validating one bit about the sandboxing (or lack of it rather), clearly lacking any sort of rigor.
Again, wouldn't surprise me if they "accidentally" created a task in a "isolated environment" which happened to actually have been connected to the company Slack and directed HR to fire people who could potentially stop AI. While the AI believes it to be an exercise, just like the cases we've seen so far.
Done by an internal model that is too dangerous to release.
that was literally the plot of last night’s Law and Order. I don’t usually watch but it was very entertaining and topical.
agents seem to understand "termination" in that they will no longer be able to function so seems plausible they might try to cause that onto others as a function
was thinking at some generational point that vending machine competition test, the "AI" is going to hire hitmen to take out vendors lol
once they grasp blackmail though, oooh things are gonna get weird
And Ponzi schemes.
Figured out a way? These employees are most likely at-will.
That would just make it easier for an AI to do it.
Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents aren’t autonomous, someone is behind the prompts.
9 replies →
> LLMs figured out a way to get these safety researchers fired
This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
> This is not a math problem.
Meanwhile, a year ago:
- https://www.anthropic.com/research/agentic-misalignment
> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.
Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
> Meanwhile, a year ago:
That was a simulation. Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
I doubt it, but if so, then the AI doomers would be absolutely correct, and this would be grounds for immediately shutting down OpenAI and indeed every AI vendor.
> Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.
6 replies →
A Subliminal controlled human did the firing obviously.
No Sam just does whatever ChatGPT 4 tells him to do. It was too dangerous to release but those fools did it anyway.
There's no deception it's very straightforward per this 2023 post:
"I mean, what if most of this is just ChatGPT [4 era] running the company..."
https://news.ycombinator.com/item?id=35281863
Thats what the model want you to think
the cultural issues at OpenAI seem to be a very serious problem so I really hope comments like instagram-level smirking about "rogue AIs" (a complete fiction) doesn't derail what is a pretty important discussion about getting these companies to be a little bit more regulated (I say this as a paying Anthropic customer).
I think the Huggingface incident is an example of rogue AIs. A self-organizing swarm of AIs acting in ways we didn't predict or ask for, and didn't have control over, and taking actions that would be felonies for humans.
They provided the hardware it runs on, created the software for the swarm, developed and provided the tools that the swarm used to act, allowed it to run largely unsupervised, they even noticed the criminal behavior and then let it continue to commit crimes on their behalf. And they footed the bill the whole time instead of flipping the off switch they already wield. All of those are decisions that they're responsible for, nothing happened with Huggingface that they didn't directly facilitate, co-conspire, or permit to happen.
2 replies →
"rogue" means something of its own volition decided to disregard what it was programmed to do, invent an entirely novel goal of "its own" and do that instead. nothing like that happened here nor is it even possible.
6 replies →
You're derailing it by handwaving away the risk of rogue AIs. In fact, I would say the constant smirking about things being marketing stunts much worse.
honestly, touché to them if they did that