← Back to context

Comment by thisgoodlife

13 hours ago

I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?

I use them. Daily. Gemini hasn’t been a contender by comparison for a long time.

  • I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do.

    Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.

    I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.

  • True but it has a niche in SQL reviews for me. Looks like Google has a lot of good sql in their corpus and in their RL digital lobotomy factory.

  • I did a test involving implementing cobol control flow in Java for a source to source translation project. Gemini was the only model to get the edge cases. Cobol is very peculiar in this regard.

    • It's very good at Elixir in my experience too. And it just does what I ask and doesn't wind me up like Opus. I don't think I've had to insult it more than once per day.

  • I had a typical $20 Gemini plan that I just downgraded to their $5 plan (to keep access to some of the models). It had been so long since I let Gemini work on (or review) any code / design / html (anything) that I couldn't justify bothering to keep wasting money on it. It fell behind badly over the past year. Astra might as well be an alien super intelligence at code compared to Gemini. I enjoy talking to Gemini, it is very good at conversation, I get solid answers to everyday questions. I intend to keep the $5 plan indefinitely for basic use. I don't expect they'll ever resurface as a competitor in coding with Astra & Fable et al.