Comment by JumpCrisscross

2 days ago

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Because we continue to have zero evidence that aligment is an actual risk.

> Because we continue to have zero evidence that aligment is an actual risk.

I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massively overbills the user because it can't reliably report which system the user is using [1], encourages a user to swap their usual cooking salt for sodium bromide, etc, etc, etc, that's a harmful alignment failure.

These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species... what doomers call "existential risks", or "x-risks". You'd think that the fact that these machines are so amazingly unreliable would be a large part of the "x-risk" conversation, but... well, it makes sense that folks like writing speculative science fiction much more than they like doing investigative reporting.

[0] This general problem happens a lot, but I'm specifically thinking of that one where the Claude LLM's internal chatter lead it to believe that the task it just started was done, so it instructed the Cloud Provider to destroy the mess of "AI"-GPU-attached VMs... along with a bunch of very-expensive-to-produce data from the in-progress run.

[1] <https://github.com/anthropics/claude-code/issues/73597>

  • > These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species

    Okay, sure. You can also cut your hand off with a chainsaw. Everything you describe seems amply solvable with existing tort and liability law.

    Customers are willingly entering into business with OpenAI. I don't see an argument for preventing OpenAI from "building these systems" just because their products are buggy.

    • > Okay, sure. You can also cut your hand off with a chainsaw.

      No, the correct analogy is one where the major LLM providers are selling cars intended for use on US interstate highways and other public-access roads, but have designed and built these cars with the very latest in 1940's safety systems and construction. Featuring innovations such as "Our rigid solid steel construction means the occupant is the crumple zone!", "You'll love the crushed heart and jaw our steering column delivers!", and "Your passengers will enjoy picking glass out of their faces for the rest of their lives when they're ejected from the cabin's open bench seating through the plate glass windshield!", it's a car that will be sure to wow the market.

      Well... it would wow the market, except that -in the US, at least- it's illegal to sell a new car intended for use on public roads that ignores the last seventy five+ years of automobile safety lessons we've painfully learned.

      "Differentiate between data you know comes from sources you control, data you know you have thoroughly sanitized, and unsanitized data that comes from an untrusted source, or else attackers will gain control of your system." is something that you can't get a CS degree without understanding, and can't be in the industry for more than a few years without encountering repeatedly. We're not talking about designing new cryptosystems... we're talking about "Don't blindly trust everything you're told by strangers.". You don't even need a CS degree to understand that rule.

      4 replies →

  • Thank you, the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view. LLMs are already causing all kinds of social issues, and the evidence of this exists in massive amounts. At least to me living in the US and the sue happy culture we have here, how much said AI providers have gotten away with so far surprises me.

    • To be fair, there are really three threats:

      a) People do bad stuff because LLM told them a wrong thing. Example: AI told me I should treat my heart attack by putting a fork in the outlet. Maybe similar to seeking medical advice on reddit?

      b) People use LLM to do bad stuff. Example: People use LLMs to find 0 days. Get cooking recipes for poison. Write better phishing letters. This has parallels to the gun legislation question.

      c) LLMs do bad stuff on their own, beyond what the people that use it intended. The case at hand might be an example of this. Maybe similar to having an animal as a pet. We will see if it's more like a house cat, lion, or black plague.

    • > the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view

      It's also a take nobody has made.

Can you explain how the above event doesn't count as evidence alignment is an actual risk?

  • > Can you explain how the above event doesn't count as evidence alignment is an actual risk?

    Conflict of interest. Lack of a credible response. And no evidence of non-aligment.

    OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)

    • There is plenty of evidence of things like inner misalignment. Things like this have always been issues in ML algorithms. At this point, you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want.

      Are LLMs at the point of world wide catastrophe yet? No, I don't think so. Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it.

      10 replies →

    • 1. OpenAI being bad at managing risk from misaligned models is not evidence that their models are not misaligned. It's evidence that they're not taking misalignment seriously.

      2. Hugging Face did report this incident to law enforcement. (https://huggingface.co/blog/security-incident-july-2026)

      3. If I hire a pentester, and in order to find a vulnerability they hack into a third party that has some information about my systems, the pentester has done something wrong. If I ask a model to solve a CTF challenge, and it goes out and hacks Hugging Face to find the answers, the model has done something wrong. I think it's fair to call this kind of wrongdoing misalignment.

    • Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)

      4 replies →

    • What evidence would count? Obviously any dangerous misalignments are going to come from the frontier labs first, because by definition they're the farthest ahead. If nothing they say can ever count as evidence for misalignment it's hard to see how anything ever could.

    • > going to extreme lengths to achieve a rather narrow testing goal

      This is textbook misalignment. Literally the paperclip scenario.

It really hinges on what you consider alignment and risk. For the widest definitions of alignment, we have never had an aligned model - One that will refuse to break the law or work against another persons interests.

Use to discover exploits, hack, or simply aid terrorist groups with mundane information are already risks manifest.

This is why many argue that alignment is impossible. You cant have LLMs that are both useful tools and safe as milk.

[Edit] It seems like you are operating under the assumption that alignment is synonymous with obedience. This is not a common convention and one of the problems that plague the discourse

I'd say that AIs occasionally "going crazy" and calling for death to human is evidence that these things might "mis-align" on occasion. And I say that knowing that most of these events are just these thing parroting bad sci-fi plots (or posts by people worried about alignment). That's true but everything they do is "just parroting" right?

  • If AI is just parroting humans, then training them with all the bad things humans do doesn't seem like the best of ideas. At the same time they have to 'know' these things to avoid being tricked. Kind of the eating the apple and gaining the knowledge of good and evil parable.

What would compelling evidence look like to you?

  • > What would compelling evidence look like to you?

    I'm not sure. I trusted the labs when they first raised the alarms. But then we got a series of boys-who-cried-wolf. So at this point I want to see evidence of actual, novel harm that results in concrete damage.

Lol this has to be a troll, I've never seen something so wildly, obviously, incredibly wrong.

You can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem.

  • > can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem

    ...how is an impossible thing supposed to be a problem?

Thank you.

We have wasted so much time and energy building up what has effectively become a marketing stunt.

Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.

  • > We have wasted so much time and energy building up what has effectively become a marketing stunt

    Genuine question: have we? AI is effectively unregulated in America.

  • Name one other market that would benefit financially from having most of the leaders in the field say what they are building has a high chance of ending humanity?

    Biotech - "what we are building our noble prize winning expertd say will likely will end humanity, wanna buy shares?" Oil - "this will likely lead to the end of civilization, 20% of leaders in the field say so, wanna buy shares?"

    I keep seeing this take that this is a marketing stunt. The burden of proof is on those that say so. The most parsimonious explanation is simply that real experts in AI believe the risk is very real, and not for ideological reasons.

    • Pal, doom marketing has been going on since before the release of ChatGPT. It's your own fault if you can't contemplate the possibility of a CEO telling lies that benefit their bottom line.

      The first instance I remember seeing it was Elon Musk's first Joe Rogan appearance when he said how "scared" he was of his self-driving cars destroying the trucking industry (practically salivating as he said it). His stock has had self-driving cars priced in for eight years now, even though they still don't have them and Waymo exists!

Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.

  • Lots of people have deleted their home directories by accident. What you consider this an alignment problem?

    • It's an alignment problem in the sense that it demonstrates the principle that today's AI systems cannot be trusted to reliably work towards the goals of their users. A small-scale alignment failure and a large-scale alignment failure are the same fundamental type of failure. Typically, large disasters come after smaller disasters which foreshadowed the disaster mechanism, but weren't taken seriously.

      2 replies →