← Back to context

Comment by carodgers

10 hours ago

This April 2026 paper is a fun and related read.

https://arxiv.org/html/2509.24239v4

Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.

> current frontier models

> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

  • So give me a falsifiable point in time, a model you would claim succeeds at the task. One does not get to hotfix-patch updater out of the pressures of reality. Today is the day.

  • The gap in capabilities is mostly quantitative and not qualitative.

    • Is it? I am on the fence on this, but it does seem like there are some qualitative improvements between the models.

      Not related to your post, but a fact I keep mulling over. The fact I don't trust the current crop of LLM's enough and I consider LLM's as a tech will hit a ceiling pretty hard, it doesn't mean parallel improvement curves won't spring up out of other research that will lead to much higher capabilities than currently.

  • The actual current frontier plays somewhere around GM level.

    https://chessbench-ai.github.io/#leaderboard

    It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.

    • I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.

      > About their ELO ratings from their own website:

      > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.

      I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..

      Please folks at least use your AIs to read stuff before making claims.

      AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.

      A GM is 2600 they can beat me in under 20 moves...

      Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.

      Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.

      42 replies →

    • Even if you take that website at face value, the ELO scores shown are relative to the other AI models tested, and not comparable to the ELO scores of humans who play against other humans.

      4 replies →

    • Probably tells us that without labs explicitly training/tuning the models or designing the harness (with fast oracle) the LLMs aren't going to get good at those areas.

    • more like you lose intelligence in chess by maxing for coding... hence knocking back the claims of emergent intelligence

    • Please do not post misinformation. They are not playing anywhere near GM level.

      "Elo is relative to the ChessBench field."

If you give the same task to an exceptionally intelligent human, who does not play chess and has only heard about it in passing, then they would be beaten by every child who has looked at the rules for more than 10 minutes.

What kind of intelligence is "playing <____> but we don't tell you the rules" supposed to test?

I suspect (in a probably ignorant fashion) that this is because learning process has been reading a lot of algebraic chess notation (such as "1. e4 e5 2. Nf3 f6 3. Nxf6 gxf6 4. Qh5! +-") then, to play, generating more of it without considering the rules of the game. This is exactly how it's always felt to me when playing chess against LLMs. Sure, "1. e4 e5 2. Nf3 Nc3" looks innocent to somebody simply learning the syntax of algebraic notation, but that Nc3 by black is an illegal move.

An LLM is the wrong approach for playing chess.

1. It’s hard to trust a 2026 paper that’s showing results for such old models.

2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.

3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

  • Good science, properly digested and presented takes time.

    The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues

  • so prove it! get a public repo out there, have it play against some open source engines

    also I think the operative letter in AGI is the G - and if the G is short for 'variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill' then its not really G at all, is it?

  • > Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

    A bash script can clone and build stockfish, feed in human moves, and reply. By your standard, this bash script would "destroy any human at chess."

    Are you interested in assessing the intelligence of the model, or the intelligence of the tools the model can use?

    • Maybe practically it doesn’t matter? Perhaps AGI is not the model but the model plus everything it’s got access to. If we’re modelling intelligence in the way we seem to have to to have any coherent definition of AGI, it seems to me <model + everything it can access> is always going to be more “intelligent” than <model> alone.

      1 reply →

  • > People who are good at it rely more on experience and deep domain expertise

    People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range.

    A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is quite low bar.

    • 1100 at online speed chess or something, could be. I'm not that deep in the chess world but everyone I know that can make 1100 in official rating can name a dozen openings and most of the known tactics, and is pretty good at applying at least one opening.

      2 replies →

  • > Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

    Delusional, but then Claude fable also isn’t beating any human at chess, the engine is.

Well, yes, PGN files have structure... But still, playing Chess with an LLM is so weird that I impulsively question the sanity of people attempting to do so. Do some people really believe training on TWIC PGNs would make an LLM a good chess player?

Llm systems are not really build for adhering to a grammar (other than "a string og tokens").

It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents.

Certainly,a harness can easily correct for it.

  • It seems absolutely crazy to me to expect an LLM to code a solution to a problem while also not expecting it to be able to adhere to a grammar.

    • How much support do we as humans need to get rules right?

      I'm an expert in my field, read my comments, my gramma is shit.

    • Why?

      You might never have tried to program before, so I don't blame it on you.

      But most programmers, even experienced ones, see grammar and type errors regularly.

  • By the promise of it, llms should be able to both adhere to grammars, or go free form where necessary. I mean, doing math is supposed to be strict but in practice it's a somewhat educated random walk in the space of correct lean theorems.

    Harnesses do correct things, sure.

    • You are right. I am imprecise.

      Languages allow a certain flexibility in their grammars - you can read a sentence without that adhering it exactly to the grammar.

      Games and programming languages (including lean) does not allow this flexibility.

      A very intelligent person would likely also reason in terms of probably outcomes before correcting a statement to adhering entirely to the grammar.

      Certainly it must be like that, otherwise reviews in math was rendered moot.

      Do we blame research mathematicians for not adhering to the grammar?

> Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

It can be very interesting and even entertaining to know where models don't do well. I don't find something like chess to be very instructive about anything though, nobody is going to pay for AI to play chess at any significant scale even if it could do it perfectly.

> The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.

I have found that not to be the case, and I am not an AI power user or cheerleader by any means. There's probably a bunch of even "simplest" tasks where AI doesn't do well and might never. That doesn't take away from the cases where it works well and is a productivity booster. It doesn't even have to be solving millennium prizes or any other breakthrough creativity or reasearch, it's still very useful in places.

This isn't how intelligence works. The LLM may not be able to play chess directly through inference, but it can write a program to do it and execute that program. Same as how human intelligence works. We can't fly, but we can build planes.

  • Human beings can play chess directly without coding up a tool.

    • Asking an LLM to play chess by writing algebraic notation is like asking a human to play chess blindfolded.

      Yes some people can do it but most people can't even if they're unusually intelligent.

      You really need to be giving the LLM a board representation.

      EDIT: I see that they actually were giving the LLMs a board representation and they still played badly. Fair enough then.

    • If they wanted to train an LLM to play chess they could easily do so.

      But nobody wants that.

The story isn't so clear cut.

The caveat is: It depends on the task.

Are there reams of chess moves that the model can train off of? No.

Are there reams of math papers the model can train off of? Yes.

  • > The caveat is: It depends on the task.

    I think the line of criticism around LLMs sucking at chess makes more sense when you understand what the AI companies are saying about the future trajectory of these models.

    The entire recursive self improvement story falls apart once you point out that there is not much "cross domain transfer learning". Meaning that training an LLM to become good at coding, math, etc, will eventually transfer into them being good at other skills that were not explicitly trained for.

    Using games like chess which have little economic value is actually a good test for this. What's even more surprising about them sucking at chess is how much information about chess strategy exists in the training data.

  • There’s multiple databases of games in algebraic notation. You can also, very easily rl train on pitting models against one another, even without mcts.

  • > Are there reams of chess moves that the model can train off of? No.

    This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.

    • It is even worse.. This is a classical reinforcement problem where data generation is easy because the rule set is pre-defined. So you really don't even need any data to start with (but would help).

      2 replies →

Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there's a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn't make any illegal moves even after a few thousand moves. Frontier labs care deeply about automating knowledge work and computer use. They are working hard on getting models better and better, and they are succeeding. Astra is a step change on that front. So good luck i guess, if chess performance is your barometer.

  • > Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player.

    If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

    • Thats just absolutly not true.

      A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.

      And there is a relevant and significant difference between the expectation of an AGI and an ASI system.

    • Pretty much this. Feed it a book or two on chess, and you should have a decent (or good) player. That's the generic intelligence people have. The aims is not to be supremely talented at something, but being able to read a manual and figure how to use/play something. Mastery can be gained overtime.

      10 replies →

    • If humans were actually intelligent, they wouldn't need to train and practice to play good chess. I mean, what level do you think people without any practice or training are ?

      11 replies →

The last post on HN I read was about someone using LLMs to reverse engineer an Apple GPU driver for linux in a month. The top comment points out how the poster must have had specialist internal domain specific contact with Apple. But then the thread concludes that wasn't the case and that this would take domain experts years to do.

> "current frontier models need laborious oversight and guardrails on even the simplest tasks"

I feel this statement is extreme. I can't personally reconcile it with any of the projects we're regularly seeing get delivered largely by LLMs now.

What are you thoughts? Like, what's your position here? Even if you sincerely believe frontier models need laborious oversight on even the simplest of tasks, do you think that accurately captures and reflects the current state and progress of frontier LLMs?

Don't get me wrong, there's lots of things LLMs can't do well, but the idea that they're basically not helpful for even the simplest of tasks seems... disingenuous?

I don't see why this is such a big deal. Nobody's using LLMs for chess, but even if they are, just give them Stockfish as part of their harness. They don't need to do everything themselves as long as they're intelligent enough to use tools.

The fact that they can play chess at all despite having no specific training for it blows my mind, and the fact it doesn’t do the same for many others shows just how far they’ve come and how fast.