Comment by p-e-w
1 day ago
> everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains
But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.
The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.
> The magnitude of improvement in unverifiable domains is small,
What makes you say that? What is an example of a domain where the improvement is small?
I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.
My argument is that if you were to compare "analyze XYZ geopolitical situation" or "explain the ramifications of XYZ law" from Opus 3.5, 4.5 and 5.5, the difference would be marginal, at least for 4.5 - 5.5. Almost all the crazy capabilities newer models have is from RLVR variants, whereas capabilities driven by RLHF are inching along.
How are the models making politics better? I don't count AI attack ads as an improvement.
2 replies →