← Back to context

Comment by TacticalCoder

3 hours ago

> I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.

I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Google/GMail) for anything that is not coding.

I find Gemini better/quicker/more polished for basically every single subject out there that is not "write me lines of code".

Same here.

I use gemini for everything not-coding, from doing research, to have custom personas for more niche topics (and feeding more detailed knowledge in these cases).

For coding and image editing, right now I find ChatGPT superior. And for software architecture designs or planning Claude is the best since a while. I still have to try Grok to be fair.

To be fair, it has been about six months since I tried a paid Google model. I will give it another go to compare. Hallucinations were the main issue back then but perhaps it has come a long way.

  • Much respect for being able to admit that you haven’t used a model in a while after saying folks probably haven’t used the models you use.

    • Okay I just tested 3.8 Flash (High) on some real world problems I have used Opus (High) and Sol (high) to solve. The outcome here is terrible.

      One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc.

      3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info.

      I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy.

      But don't take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt...)

      This has been a useful exercise. It's important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.