Comment by shaky-carrousel

1 year ago

That's easy to test, invent a new chess variant and see how the model does.

20 comments

shaky-carrousel

You're imagining LLMs don't just regurgitate and recombine things they already know from things they have seen before. A new variant would not be in the dataset so would not be understood. In fact this is quite a good way to show LLMs are NOT thinking or understanding anything in the way we understand it.

shaky-carrousel 1 year ago
Yes, that's how you can really tell if the model is doing real thinking and not recombinating things. If it can correctly play a novel game, then it's doing more than that.
- jahnu 1 year ago
  
  I wonder what the minimal amount of change qualifies as novel?
  "Chess but white and black swap their knights" for example?
  
  2 replies →
- timdiggerm 1 year ago
  
  By that standard (and it is a good standard), none of these "AI" things are doing any thinking
  
  3 replies →
- dwighttk 1 year ago
  
  No LLM model is doing any thinking.
  
  4 replies →
empath75 1 year ago

You say this quite confidently, but LLMs do generalize somewhat.

dmurray 1 year ago

In both scenarios it would perform poorly on that.

If the chess specialization was done through reinforcement learning, that's not going to transfer to your new variant, any more than access to Stockfish would help it.

gliptic 1 year ago

Both an LLM and Stockfish would fail that test.

delusional 1 year ago
Nobody is claiming that Stockfish is learning generalizable concepts that can one day meaningfully replace people in value creating work.
- droopyEyelids 1 year ago
  
  The point was such a question could not be used to tell whether the llm was calling a chess engine
  
  1 reply →