← Back to context

Comment by ricardobeat

11 hours ago

For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great.

Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).

https://deepswe.datacurve.ai/blog/deepswe-v1-1

gemini 3.8 flash is also 74% and google just started letting all their engineers use claude...go figure

  • > and google just started letting all their engineers use claude

    That's misleading.

    1. Having different models available is useful for A/B testing and helping improve Gemini itself.

    2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.

2.6-pro just reached 63.7% by step 10, it's on step 11 right now.

Even flash reached 60.7% by step 12, and it's on step 16 now.

This is so exciting lmao.

  • DeepSWE is saturated now IMO, and is basically worthless. Lots of new models get around 74%. Shame too, because it was a pretty decent benchmark for a few months there.