Comment by supermdguy
10 hours ago
Here's the performance of frontier models without reasoning, to more directly address the claim that raw performance is plateauing:
https://artificialanalysis.ai/evaluations/artificial-analysi...
I don't have any insider info, but if model sizes actually have increased exponentially since GPT 4.1, there's an argument to be made that there are diminishing returns in scaling pretraining alone.
Also interesting thing I haven't noticed before, Opus models have followed a really consistent linear improvement, while it looks like OpenAI struggled with base model performance until 5.5/5.6 (EDIT - 5.5 was their first new pretraining run in over a year).
The trend I've found most interesting is models of the same size getting better.
I'm very much looking forward to seeing how Qwen 3.8 27B compares to Qwen 3.6 27B next week, for example.
And the latest DeepSeek v4 Flash has extremely impressive performance for a 304B model.
The trends you found don't support my goals so I've got some other trends I find more interesting than yours.
What are my goals here?