← Back to context

Comment by bertili

10 hours ago

DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!

when are we going to stop pretending these benchmarks have any meaning?

anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

  • +1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.

    (I'm not happy about the above being true, but it's the reality I seem to inhabit.)

    • And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.

    • I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo

With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.

  • and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.

    This technology is strictly an extractive parasite on the world. Use it, but don't be excited.

Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.

Compare that to Muse spark 1.3

$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)

It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.

But is the score really reflective of the quality or are both models benchmaxxing?

  • Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.

  • Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.

    This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.