Comment by dudeinhawaii

4 days ago

Grok is quite interesting. I run comparisons almost daily on tasks and Grok is its own beast, in a good way.

It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same thing with diminishing quality. It makes the lineup pointless. Ditto for Anthropic. Gemini-3.6-Flash and 3.1 Pro genuinely behave differently. Opus 5 and Fable are.. cousins.

I find that when I want to test a complex creative challenge, having 4 "families" to choose from makes the experience interesting since they will excel in different areas.

Grok might implement unique lighting, Opus, elegant primitives, Sol, accurate snowfall in one pass, Gemini, silky movement. Combined, you can pick and choose best.

For what its worth, Grok always feels "messy" but finishes. Grok 4.6 though is no longer "smart and fast". It's about as fast as Sol though.

A big improvement I noticed in 4.6 was tool use for verification. Previously, Opus/Fable were the only models to consistently screenshot things that they can't directly interact with easily. Now Grok is probably right behind them, perhaps tied with Sol on propensity to verify visually. Grok 4.5 notably did not do this often.