Comment by scotty79
3 days ago
On DeepSwe it's strictly beaten by Luna on max, cost and result.
Damn, Luna on max is as good on DeepSWE as Kimi k3, I think I dismissed this model unjustly.
3 days ago
On DeepSwe it's strictly beaten by Luna on max, cost and result.
Damn, Luna on max is as good on DeepSWE as Kimi k3, I think I dismissed this model unjustly.
This is why benchmarks are scary, since Artificial Analysis puts it a fair bit behind Kimi K3.
Kimi K3 is a beast though, just costly.
Still cheaper than API rates for Opus!
So in a way Luna is the new Gemini Flash? I've been out of the game for a while.
Luna is the smallest variant of gpt-5.6 from openai and it seems to beat Gemini Flash 3.7 on DeepSWE benchmark (which is one of the new coding benchmarks that people find more relevant to their daily work than old, saturated and gambled benchmarks). It beats it massively on cost and by a bit on quality.