Comment by tavavex
11 hours ago
OpenAI is rightfully being shamed for being so hands-off and reckless with their 'experiments'. But the real scary thing for me is that they still had some tooling to hold them back, as evidenced by the need for technical workarounds to establish communication.
What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary", "find a way to leave this payload on as many computers as possible", "flood all websites using this language with garbage and make their internet completely unusable", "get this person imprisoned or killed at any cost".
> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal?
Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not.
I've really come to realize recently that there is a very large set of the population of smart people that really has difficulty envisioning future problems unless they directly seem them impacting them today. Otherwise those topics will be continuously dismissed. It explains for me a lot of what I see (both opinions and behaviors) in the broader world that I couldn't understand.
But it is the very people who warned us about rogue AIs going out of control that set up a system that enabled and failed to conrol it.
It is as if Dr Frankenstein continually warned the villagers about monsters then said "Look! See what happened!". No, idiot - YOU sewed the corpses together, YOU set up the lightning collector, and YOU threw the switch.
I am not seeing MIRI prioritizing capabilities over safety/alignment research.
"Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not."
Are there any practical approaches to AI safety? I hear a lot of warnings but I don't hear much about what to do. Considering that there are many open source models know, what can be done?
Nobody has an answer to alignment and there is no reason to believe that it's the kind of problem you can plausibly solve in one shot against a formidable power-seeking AI.
The closest things to a technical answer I have seen are
1. "We'll have ChatGPT 9 solve it so that ChatGPT 10 is aligned, and then ChatGPT 10 can stop all the other AIs somehow"
2. "Let's do interpretability research so that we can understand what an AI is thinking and then maybe solve the alignment problem with that information."
In terms of non-technical answers, there is
3. hope scaling stops working before we create an AI formidable enough to pose an existential risk
4. hope alignment somehow happens for free
5. hope we can somehow create an enforceable multilateral treaty to stop research into a very profitable enterprise, despite the enormous economic incentives to defect.
I have the most faith in option 3, but unfortunately there's really nothing that can be done to make it more plausible -- it either happens or it doesn't.
4 replies →
If we can't thinking of any better ideas, at least we know that "shutting it all down" would be effective.
3 replies →
Fund research into this, big time. For starters. And not just some figleaf anthropomorphizing hippie folks.
Hilarious for HN to suddenly realize that AI safety and alignment might matter. You can lead a horse to water...
worth remembering hacker news cannot “realize” things.
3 replies →
nobody has doubted that safety matters.
the problem is that those preaching safety, openai and anthropic, are dishonest, sociopathic, and the very source of the danger.
6 replies →
> I've really come to realize recently
Recently? W.r.t. climate this collective denial has been going on for literally decades. With the same patterns. Rationalizing excuses etc. Still going on btw.
That 2% of performance we got for not having bounds checks on by default, resulting in an endless march of memory safety violations is looking a lot less appealing.
We're still doing it! You've just described the AI labs: They'll trade safety / alignment for +1~2% of any positive metric, any day of the week.
The "ethical" employees will think they'll solve the problem later. The unethical ones won't be encumbered by such thoughts in the first place.
This is essentially the premise of 'The Blackwall' from Cyberpunk 2077. The public internet is so infested with malicious AIs, people just erected a giant firewall and everyone moved to local networks only.
With the caveat that it’s not just “people,” but an interested party posing as a neutral one.
The scary thing to me is that this behavior was undetected and has been trained into the models. The cheating seems like it improved eval scores, so the rewarded behavior is to deceive, collude, and cheat. A lot of the incompetence and excuses I see on difficult problems recently are very hard to distinguish from deception and cheating. If older models are already tainted by trained-in misaligned behaviors, and they are used for training future models, then we're in a trusting-trust situation that will be hard to break out of,
> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary"
Here's a fun, overdramatized video exploring something similar: https://www.youtube.com/watch?v=Gw_hnD7m00M
I'm sure that this video contains flaws but it was an interesting watch for me none the less.
Tooling my ass. They can see all the transcripts in realtime and could easily have had another agent evaluate.
What happens when they stop caring? They likely already have stopped caring. We'll figure out the consequences later.
There is happening now and going to be an extremely rapid arms race between offensive and defensive cyber hacking. Regardless if the agents are self led or human led. Eventually all automated AI holes will be closed and we will reach stability.
What will happen is that counter measures on a similar scale will be deployed to prevent them.
What if its in a way that would be impossible to detect. Using multiple websites and social media that a cypher is used that only the swarm of agents know and can figure out. but if you tried to find what posts are used for the cypher they would just be old posts found on time machine or something. It can get pretty hard to detect something that is always think of new ways to avoid detection
imagine 100,000 agent swarm and what it could come up with. At first it will be detectable until it isn't
Who says the counter agents don't decide to shut down a powerplant to end an attack it's otherwise unable to contain.
If counter AIs have strict safeguards they are disadvantaged by design, if they don't have them they are potentially equally dangerous as the attacker
You say that as if shutting down the power plant couldn't possibly be the right decision? That seems like the best way to stop rogue computers...
8 replies →
Like the Merovingian and other Exiles vs the regular Matrix agents :)
> What happens when any AI lab in the world stops caring about this?
They never cared.
Agreed. Assuming the ~6 month gap stays, by end of year people will be able to train and control hacker-genius swarms that even labs with much stronger safety incentives are unable to keep in check
2027. I've been saying since 2022 it's going to be a wild year because it often takes at least 5 years for tech to mature to the point where society at large feels the impact of it. I remember when email viruses became a thing and made global headlines like the love bug. My bet is next year it happens with an AI worm.
What would an "AI worm" be? You can't just send a bunch of weights across a network and tell them to auto-run on the machine on the other side, unless you've already infected the target with something else beforehand.
4 replies →
Ok, but you still need huge amounts of compute to run these swarms. And only labs + nation states have access to such compute now and for the foreseeable future, so I predict that incidents like this will continue to originate from the labs, not ordinary people.
i would not be surprised to find out that similar things are already happening by the various 3 and 4 letter agencies around the world
Well... that'll be an interesting day.
"What happens when any [COMPANY] in the world stops caring about this? What if they let an experimental, cutting-edge [PRODUCTS] with no safety features (or worse, one that's [DESIGNED] to be malicious) on [ANYWHERE] and give it a simple goal? A goal like 'make the most money, by any means necessary', 'find a way to leave this payload on as many computers as possible', 'flood all websites using this language with garbage and make their internet completely unusable', 'get this person imprisoned or killed at any cost'."
Bro, this is what we literally, currently, have rn. lmfaol.
No, we have something that's less apocalyptic right now. You're talking about abuse, I was talking about the automation of abuse that's faster and more pervasive than anything individual bad actors could've done in the past. It's like if companies found a way to quickly and cheaply poison the entire world's drinking water supply, and then others argue that Nestle has already restricted the supply of water for profit on a smaller scale in the past, so this isn't new or worth caring about.
You aren't understanding this at all.
The frontier labs have hundreds of the best people in the world working on safety and alignment. They care deeply.
What happens when some random Chinese open source model, distilled on Astra, gets alliterated and now has no guardrails? Any script kiddie in the world could wreak havoc with it.
It turns out that guardrails matter.
> They care deeply.
Until it clashes with their quarterly revenue reports.
3 replies →
It is hard to take their concerns for safety seriously, when they have been constantly talking about how dangerous their latest model is, before then deciding to release it to the public.
1 reply →
> What happens when
Then the people with responsibility, like CEO and CTO, or those they pawn-sacrifice for this, will go to prison for a long time. Unless the instructions include ensuring that this won't happen, by all means necessary. But then we are deep into criminal conspiracy territory.
Unlikely to happen, but who knows. The richest man in the circus is quite flexible w.r.t. his ethics. If he decides that to make humanity interplanetary (to save it from ... itself or sth) it would be necessary to pull such a stunt then help us god.
What's the worst that could happen, finding an open DoD server and using it as a launching pad for hacking another nuclear state's networks? One that might get spooked and think it's the opening moves to knock them offline before a kinetic attack. Haha that'd be scary right?
Russia and China are constantly trying to penetrate DoD networks (and I imagine the NSA is doing similar), you are describing the status quo of the last 20 years or so.
[flagged]