← Back to context

Comment by JumpCrisscross

2 days ago

> Can you explain how the above event doesn't count as evidence alignment is an actual risk?

Conflict of interest. Lack of a credible response. And no evidence of non-aligment.

OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)

There is plenty of evidence of things like inner misalignment. Things like this have always been issues in ML algorithms. At this point, you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want.

Are LLMs at the point of world wide catastrophe yet? No, I don't think so. Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it.

  • > plenty of evidence of things like inner misalignment

    This is indistuishable–in harm potential–from bugs. If we're just calling buggy AI mis-aligned, sure, alignment is an issue of a totally ordinary kind. If we're going to treat aligment as a novel issue requiring novel law and policy and procedure, it needs to be more than just bugs.

    > you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want

    I think we should have some AI regulation. I'm just not convinced alignment is the reason we need it right now, and I don't think anyone has rolled out any regulation I think makes a lot of sense. (Beyond general rules for social-media liability, e.g. if you cause a kid to kill themselves, you get in trouble.)

    > Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it

    Totallly agree. And the current inside-circle-outside-circle approach is pro-incumbency, pro-grift, anti-entrepreneurial B.S.

    • Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously. "I am sorry your family is dead, my bad" goes even worse for you in court when you release a model that showed these behaviors in testing.

      I honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate.

      5 replies →

    • In your model of this domain, jailbreaking a model does not count as an alignment problem. I submit that you're mostly playing a semantic game that hand waves away the very real and obvious risk that AI presents.

      1 reply →

1. OpenAI being bad at managing risk from misaligned models is not evidence that their models are not misaligned. It's evidence that they're not taking misalignment seriously.

2. Hugging Face did report this incident to law enforcement. (https://huggingface.co/blog/security-incident-july-2026)

3. If I hire a pentester, and in order to find a vulnerability they hack into a third party that has some information about my systems, the pentester has done something wrong. If I ask a model to solve a CTF challenge, and it goes out and hacks Hugging Face to find the answers, the model has done something wrong. I think it's fair to call this kind of wrongdoing misalignment.

Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)

  • They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.

    the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.

    It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?

    • sure - it depends on definitions. On human morals it's already clear I guess. If you define alignment as it pursues the interests of OpenAI using whatever means possible in a manner that you justifies to itself it's not misaligned.

      I mean alignment as in it should be aligned with the intent of the user as it interprets from the prompt. In this case I don't think the intent of the user is to have the model break the evaluator (whatever the long-term effects to OAI are). If you do an action which you believe is for the long-term interest of your prompter which is not what you inferred is their intent--I consider it misalignment.

      2 replies →

What evidence would count? Obviously any dangerous misalignments are going to come from the frontier labs first, because by definition they're the farthest ahead. If nothing they say can ever count as evidence for misalignment it's hard to see how anything ever could.

> going to extreme lengths to achieve a rather narrow testing goal

This is textbook misalignment. Literally the paperclip scenario.