← Back to context

Comment by minraws

12 hours ago

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.

> About their ELO ratings from their own website:

> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.

I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..

Please folks at least use your AIs to read stuff before making claims.

AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.

A GM is 2600 they can beat me in under 20 moves...

Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.

Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.

>> I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.

This is unfair to HN readers all of whom but one did not post the comment you replied to. You can't just tar everyone with the same brush. There are thousands (hundreds of thousands?) of users on this site.

  • How many posts if I link that do the same thing will you agree this is the norm here.

    Not everything I have the time and energy to reply to. This chess one is just ridiculous claims on top of ridiculous claims all the way and 0 push back in the comments except mine.

    I don't even know if there is critical thought or we believe what we read/shared/etc

    • No, I don't agree it's the norm. There is though a general tendency to opine with strong views on subjects posters have no expertise on. I think that's because many are software engineers (or equivalent) and they are used to being expected to "wing it" on whatever technical subject comes up. On the other hand you can always find informed comments by users who have specialist knowledge.

      And there's plenty of pushback on here about the chess thing besides your very valid points.

      EDIT: anyway if I can offer a bit of unsolicited advice, it won't do you or anyone any good to accuse everyone who doesn't agree with you of laziness, even if you can see e.g. they haven't really read an article. Just say the thing you wan to say and let them figure it out. Most people will appreciate that much better and you will feel better about yourself for acting like a mature adult.

      It's even in the site guidelines:

      Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".

HN is no different than Reddit, or any social media for that matter, in that commenters pretend to read articles.

  • Back in 2001, our social medium was Slashdot and no one ever pretended to read the article. No one read the article either. It was slashdotted most of the time anyways.

  • that is if it even a human commenter at all

    • State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behavior in these systems that's difficult to engineer out.

      5 replies →

What levels are they actually at in your experience?

  • Sub 1300 that's my rating in the singular official tournament I participated at.

    But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves).

    I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win.

    I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it.

    If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800.

    700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.

    • As someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself.

      Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win.

      Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous

  • So you can see an actual game on that website, and the play seems pretty decent to me for a while (~1700 lichess = 1300 elo) until move 28 when black throws away their queen for absolutely no reason in an incomprehensible blunder.

    In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.

    • I've always liked the analogy that talking to an LLM is like talking to a really, really smart person with a head injury.

> even if I give them literal infinite time and all the subagents and internet access..

Don't use the word infinite in any CS claims. They can recreate or approximate monte Carlo tree search and it technically is still a correct solution in your framing of the problem so long they defeat you.

  • If they don't actually do that, given "infinite" time, then it doesn't matter what they allegedly "can" do.

The AI can write a chess bot program that will beat you.

You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.

  • > The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

    This argument is fundamentally incompatible with all the breathless rhetoric about "AGI" coming from the providers' general direction.

  • I can write a chess bot program that will beat you. Does that mean I’m good at chess?

    >If they cared to have it perform well in chess games, you'd see a different shape and behavior.

    So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?

  • So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here.

    Delusion runs deep in HN circles.

    I say that as someone heavily invested in AI startups and projects and as someone working in the field.

    I think most people on HN should touch grass and find real human contact. Lmao

    Incredible reasoning all around here.

    • An AGI doesn't stand for 'perfect intelligence' it stands for artificial general intelligence.

      And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.

      Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.

      This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing

      2 replies →

    • AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI

      Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

      23 replies →

    • I'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well.

      People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.

      8 replies →

  • First, you're moving the goalposts. Second, it's not actually true that any existing frontier AI can write a chess bot program that can beat a 1600 player ... not unless the program is derived from Stockfish or some other leading engine that has been in development for decades.

    > The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

    These comments indicate a complete failure to understand the technology.

    I won't respond again.

> I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access.

I don't believe this.

You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.

A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.

I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.