Comment by firasd
10 hours ago
Opus 4.8 (max thinking) scored highest and Grok 4.3 lowest
It's hard to understand what's going on with Grok. It's like it has capabilities in a theoretical sense but maybe the training is so focused on being in x.com/grok.com with the web search tool enabled for "is this true?11" type queries that with any API type usage with document workflow instructions, tool use, code gen etc it completely falls over
After they acquired Cursor, Grok 4.5 seems like a completely new model, performing at Opus 4.6 level, I'd say. But much cheaper and faster.
maybe it is Grokimi?
https://venturebeat.com/technology/cursors-composer-2-was-se...