Comment by slowin
16 hours ago
Really? I've used all of the models extensively and Grok doesn't even compare to the others. I can ask for a big task and it will say "Done!" like 3 minutes in, but it will have done just the least amount of work possible. I also notice "muskisms" leaking back in the results like "I'm not roasting your code". Completely unusable for me and not comparable to the other model providers.
Interesting, havn't noticed that. Mine will work for hours. Maybe i'm still in the "get them hooked" phase. Just started using it last week.
That is interesting, I've never been able to get it to successfully complete big tasks to any level of quality. For me the Anthropic models are by far the best at that, but the newer OpenAI models are starting to approach that quality. Grok always seemed like it was designed to write nasty tweets and the code quality I got from it matched that.
I'm considering switching as well, used the API to code with Grok via pi, was really pleasant whereas using Claude just gives me a cortisol spike
The only model I've found to beat it for coding (Python, Rust, Kotlin, PHP) and planning is Fable. What are you using it for?
I'm using it for coding and it's never approached the level of Opus for me. Sol is pretty good too.
I actually haven't used Sol yet. I'd love to hear what you use it for, and how it performs.
1 reply →
When did you try it? Grok before 4.5 was way worse in my experience.
I don't remember the version exactly but maybe June/July timeframe? I tried it on maybe a half dozen projects with different types of asks as I was comparing models for my company at the time.
Grok 4.5 came out in mid July and is a huge leap forward. Not quite as good as Opus 5 for code in my experience, but better and more succinct for understanding human documents.