← Back to context

Comment by lelanthran

4 hours ago

> We're not talking about learning the rules of chess here, but playing a competent game from just being shown the rules.

Okay, lets go with that: it's the "shown the rules" bit that we are arguing about.

The argument is that a human may play maybe a dozen games after learning the rules, after which they won't be inadvertently attempting illegal moves. What we are observing with SOTA models is that, even after seeing millions of chess rules, rulebooks, actual games, etc, they still attempt illegal moves.

This does not point to generalisable and adaptable intelligence, such as we see in the average human.

This is not good reasoning. Humans need at least dozens if not hundreds of reinforcement sessions to only make legal moves, and still occasionally fail (consider pins, discovered check, failing to respond to check). LLMs must one-shot a competent game after imbibing a mass of disconnected units of information about chess. Nothing about the two are similar.

See my comment here for more: https://news.ycombinator.com/item?id=49725306