Comment by Gareth321

2 hours ago

Okay I just tested 3.8 Flash (High) on some real world problems I have used Opus (High) and Sol (high) to solve. The outcome here is terrible.

One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc.

3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info.

I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy.

But don't take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt...)

This has been a useful exercise. It's important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.

Gemini's search harness in the Google app is (ironically) bad so it makes the model look bad.

If you really want to compare apples to apples you need to test Gemini models against other models using the same third party search harness.

Otherwise you are largely measuring how much computation the model provider is allocating to a search harness.