← Back to context

Comment by doctoboggan

8 hours ago

Someone should make the millennium problems for AI alignment so the frontier labs will actually try to solve that problem.

> Someone should make the millennium problems for AI alignment so the frontier labs will actually try to solve that problem.

Rather: Someone should make the millennium problems for human alignment so the top researchers will actually try to solve that problem. :-)

A group already (sort-of) has for the broader theory: https://arxiv.org/abs/2604.21691

It's a good list, but "alignment" is much more targetted. I attended a workshop with researchers from OpenAI and Anthropic also present, where the objective was to hash out what the problems in alignment even are. This, it turned out, was extraordinarily difficult.

The problems with alignment are to do with how we define and search for this vague notion when it is mathematically ill-posed at present.

You'd think that believing they are going to destroy humanity would motivate them.

  • If you ignore what people say and instead observe what they do, the world makes a lot more sense.