← Back to context

Comment by user43928

2 hours ago

There is no need to ask. If you want to test SOTA models today, there are obviously only two: GPT-6 Astra and Fable 5.1.

The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier.

That Claude version is no longer available today, Gemini 2.5 Pro will be shutdown next month, and the OpenAI models are only available via the API today.

Fortunately, a fellow commenter was so kind and did it with Astra. Didn't do that well either [0]. I'm sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)...

I'll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. "Just" having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn't even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless.

[0] https://news.ycombinator.com/item?id=49720751

  • Doesn't look impressive, although I'm hearing a marked improvement in choosing legal moves, compared to early 2025.

    Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better?

    I would not be surprised if OpenAI released a model that beats humans at chess this year.