Comment by andai

16 hours ago

Wait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?

IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High.

But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny".

Flash is my go-to for prototyping, and basically anything that isn't writing production code.

  • The only company with a proper TPU set-up is bound to have the fast models, now add a market cap like Google to the mix.

  • Same. Love oneshotting or sanity checks. Which fortunately is a lot of my workflow (lot of long tail stuff fits in one prompt).

  • Luna is way slow. I don't remember an OpenAI model ever being this slow.

    edit: I have a subscription; direct call.

    • Are you using direct or via OpenRouter? I think OpenRouter Luna always uses the `flex` tier, which is quite a bit slower.

  • Its not good at not making mistakes, but what it produces is structurally quite nice, not over-engineered (looking at you Sol) and its personality isn’t annoying (looking at you Claude). A bit like Grok Code, but Grok is a better coder.

It's pretty good if you can actively steer it , its actually really really good , the antigravity free tier and pro tiers are generous as well . I'm shocked at how fast it generates tokens.

They're quite selective in benchmarks, c.f. notably only bad one is 10% on TerminalBench. It's a really addled model, one time I said "Hi" and it built out a 4 panel hello world app with (fake) weather, a todo list, and a couple other things I forgot. I wouldn't be comfortable saying "ignore the #s!" except when I complained it was trash and way overcooked on agentic coding yet not good at it, and a couple DeepMind ML people liked the tweet.

  • i discount people who lean too heavily into benchmark as the authoritative truth when it comes to evaluation of coding capability of these models.

    experience tells me that those people simply have not used models for a long period of time specifically on coding and have run their own comparisons

    to someone who uses all vendors, the differences are very palpable and drives purchase decisions.

    also keep in mind Gemini and other labs have repeatedly done benchmaxxing, you must have your own benchmarks to evaluate these models.