← Back to context

Comment by notatoad

6 hours ago

when are we going to stop pretending these benchmarks have any meaning?

anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

Yep these software benches are only good at testing how well they can one shot. For the kind of attended/assisted development most of us do with agents it’s hard to find a benchmark that reflects my own experience of the frontier models still being quite far ahead.

+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.

(I'm not happy about the above being true, but it's the reality I seem to inhabit.)

  • And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.

  • I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo