← Back to context

Comment by Sharlin

4 days ago

If you need "guardrails" to ensure (an illusion of) alignment, you’ve already lost. It’s like using a denylist to avoid SQL injection.

Then what you're saying is we should not build AI.

Simply put you cannot have generic algorithms/intelligence without the potential of 'unaligned' behavior. In humans we have all kinds of punishment systems for dealing with unaligned behavior post ad hoc because people do all kinds of unaligned stupid shit.

Making powerful AI may be one of those things that the only winning move is not to play.