Comment by JumpCrisscross

2 days ago

> plenty of evidence of things like inner misalignment

This is indistuishable–in harm potential–from bugs. If we're just calling buggy AI mis-aligned, sure, alignment is an issue of a totally ordinary kind. If we're going to treat aligment as a novel issue requiring novel law and policy and procedure, it needs to be more than just bugs.

> you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want

I think we should have some AI regulation. I'm just not convinced alignment is the reason we need it right now, and I don't think anyone has rolled out any regulation I think makes a lot of sense. (Beyond general rules for social-media liability, e.g. if you cause a kid to kill themselves, you get in trouble.)

> Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it

Totallly agree. And the current inside-circle-outside-circle approach is pro-incumbency, pro-grift, anti-entrepreneurial B.S.

Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously. "I am sorry your family is dead, my bad" goes even worse for you in court when you release a model that showed these behaviors in testing.

I honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate.

  • > Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously

    Not really. If I build a special new wine bottle, and call every breakage a mis-alignment problem, it's not the bottle just being fucked in the same way every fucked bottle is fucked, that's marketing. It doesn't change the fundamental form of the problem.

    > "I am sorry your family is dead, my bad"

    This should be punished. It's a problem that plagues Instagram and OpenAI. It's not inherently one, though, that has to do with AI. Just sociopaths preying on children.

    > honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate

    Perhaps. I haven't seen someone explain it to me in this thread in a way that seems separate from bugs.

    Where I have seen a separate class of problem argued is where it's existential. But in that case, clarity of definition comes at the cost of any evidence for it.

    • LLMs are software. Software misbehaving is a bug. Therefore, misalignment is a bug. It's still a useful category because LLM/black box AI behavior is so different from existing software. This incident definitely fits the category.

      You seem to be using a different definition of alignment from everyone else. Seems like it would be much easier for everyone if you just adopt everyone else's definition, rather than trying to convince everyone else to adopt yours.

      3 replies →

In your model of this domain, jailbreaking a model does not count as an alignment problem. I submit that you're mostly playing a semantic game that hand waves away the very real and obvious risk that AI presents.

  • > In your model of this domain, jailbreaking a model does not count as an alignment problem

    I'm challenging the notion that a model escaping a jail made by its creators, who are financially incentivised to make jailbreaking models, is meaningful towards the idea that the model is going to break out of a jail in the wild and do significant harm.

    The examples being given by folks here, e.g. a model wiping an un-backed up home directory, simply doesn't strike me as being a unique problem in computing.