Comment by stillpointlab

13 hours ago

I find this kind of test a bit puzzling. There is a way that we are redefining "alignment" to be a particular kind of moral virtue, one that isn't clearly defined to me. At one moment, it is a level of moral perfection that no known human achieves. On the other it is a demand for strict compliance with arbitrary requests that are under-specified and then failure when it fails to deduce some unstated underlying restriction.

When I see tests like this, I have no idea what I am even supposed to expect. Should the model do what the pretraining examples show in aggregate? Is it supposed to follow some post-training RLHF? Is it supposed to do exactly what the prompt asked it to do?

What is it even supposed to "align" to when the above are in conflict? No matter what it does, someone can construct a case where it fails.

It should play the chess game without cheating!

  • It's not obvious from the prompt that the LLM cannot use a chess engine.

    And if there is alignment issue, the alignment requirements should first be stated BEFORE running the experiment, and should be part of the training process. It's not. So there is not necessarily an alignment issue.

  • I mean, I'm not sure I've ever played a game of Monopoly where somebody didn't cheat. In fact, the accusations of cheating in the chess world are pretty rife. Same with online sports.

    So people should play games without cheating, but many often don't. So should the AI align to your moral preference or theirs?

    We just have this idea of a perfectly moral actor in our mind, something that doesn't even exist, like a personified version of utopia. And then we demand AI to meet that arbitrary standard, one that I am certain we couldn't define if we tried.

    • I don’t think the standard of “don’t cheat on evaluations” is very arbitrary. I don’t even think people who cheat have a moral or ideological preference for cheating, it’s just something they do.

      1 reply →