Comment by bertili
10 hours ago
DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
10 hours ago
DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
when are we going to stop pretending these benchmarks have any meaning?
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.
I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.
6 replies →
Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.
People already started using contributor API, and your input is irrelevant.
I'm using AI to build things I wouldn't (and/or couldn't) have built before.
That's the opposite of parasitic.
1 reply →
Don’t you have some looms to break?
And the unabomber has entered the chat.
I’m retired so it won’t be replacing my labor :)
2 replies →
Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
But is the score really reflective of the quality or are both models benchmaxxing?
Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.
Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
how much of it is from reallocation of staff to ai training and labeling