← Back to context

Comment by jdthedisciple

6 hours ago

I dare anyone to convince me the benchmarks are not meaningless.

Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?

How would this alleged difference (most likely bs) actually show up in reality?

GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.

Ah, so you didn't read the article.

>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

  • You didn't read my question, bc that excerpt doesn't answer, nor do they demonstrate

    > how would this alleged difference (most likely bs) actually show up in reality?

    Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright

    All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?