Comment by joefourier

12 hours ago

> current frontier models

> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with:

> Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position.

(I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.)

For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason.

It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the "book" theory in its training data, but it still completely fell apart at early midgame.

  • So what? Give it to an agent and it will find and download the best computer chess programs available and absolutely thwomp you. I doubt you could keep a chessboard in your head with 100% precision either. Transformer models just don't have state for that.

So give me a falsifiable point in time, a model you would claim succeeds at the task. One does not get to hotfix-patch updater out of the pressures of reality. Today is the day.

  • There is no need to ask. If you want to test SOTA models today, there are obviously only two: GPT-6 Astra and Fable 5.1.

    The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier.

    That Claude version is no longer available today, Gemini 2.5 Pro will be shutdown next month, and the OpenAI models are only available via the API today.

    • Fortunately, a fellow commenter was so kind and did it with Astra. Didn't do that well either [0]. I'm sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)...

      I'll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. "Just" having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn't even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless.

      [0] https://news.ycombinator.com/item?id=49720751

      1 reply →

  • Well I guess this excuse is finally gonna become obsolete soon with all the "pacing" nonsense.

The gap in capabilities is mostly quantitative and not qualitative.

  • Is it? I am on the fence on this, but it does seem like there are some qualitative improvements between the models.

    Not related to your post, but a fact I keep mulling over. The fact I don't trust the current crop of LLM's enough and I consider LLM's as a tech will hit a ceiling pretty hard, it doesn't mean parallel improvement curves won't spring up out of other research that will lead to much higher capabilities than currently.

    • > but it does seem like there are some qualitative improvements between the models.

      It could easily seem that way, I think, in a "quantity has a quality of its own" kind of way. When you can come to the same conclusion faster, that lets you iterate more; and sometimes when you iterate you find more things.

"Transatlantic flight will never be commercially viable, we conclude based on careful study of several aircraft designs from the 1920s"

The actual current frontier plays somewhere around GM level.

https://chessbench-ai.github.io/#leaderboard

It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.

  • I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.

    > About their ELO ratings from their own website:

    > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.

    I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..

    Please folks at least use your AIs to read stuff before making claims.

    AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.

    A GM is 2600 they can beat me in under 20 moves...

    Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.

    Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.

    • >> I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.

      This is unfair to HN readers all of whom but one did not post the comment you replied to. You can't just tar everyone with the same brush. There are thousands (hundreds of thousands?) of users on this site.

      1 reply →

    • > even if I give them literal infinite time and all the subagents and internet access..

      Don't use the word infinite in any CS claims. They can recreate or approximate monte Carlo tree search and it technically is still a correct solution in your framing of the problem so long they defeat you.

      1 reply →

    • The AI can write a chess bot program that will beat you.

      You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

      We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.

      39 replies →

    • > I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access.

      I don't believe this.

      You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.

      A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.

      I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.

  • Even if you take that website at face value, the ELO scores shown are relative to the other AI models tested, and not comparable to the ELO scores of humans who play against other humans.

    • I wonder why they didn’t throw a real chess engine in there for a baseline. There are engines where you can set the elo in the settings, so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other.

      3 replies →

  • These ratings seems very wrong, i have beaten GPT Astra max thinking in chess and my rating is close to 1500. The ratings here seem more accurate: https://chessbenchllm.onrender.com/

    GPT-6 almost never suggests an illegal move anymore while even Sol still did so time to time

    • "Elo is relative to the ChessBench field."

      They are of course "wrong" if you don't read the faint fine print and sensibly interpret them as FIDE or similar ratings.

  • Probably tells us that without labs explicitly training/tuning the models or designing the harness (with fast oracle) the LLMs aren't going to get good at those areas.

  • > The actual current frontier plays somewhere around GM level.... It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors

    Sorry, but I am not buying that 5.6-Sol is that much better than 5.6-Luna, which can barely be coaxed to reach the midgame with legal moves and an apparent understanding of what the position is.

  • more like you lose intelligence in chess by maxing for coding... hence knocking back the claims of emergent intelligence

  • Please do not post misinformation. They are not playing anywhere near GM level.

    "Elo is relative to the ChessBench field."