← Back to context

Comment by CuriouslyC

20 hours ago

The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.

> The magnitude of improvement in unverifiable domains is small,

What makes you say that? What is an example of a domain where the improvement is small?

I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.

  • My argument is that if you were to compare "analyze XYZ geopolitical situation" or "explain the ramifications of XYZ law" from Opus 3.5, 4.5 and 5.5, the difference would be marginal, at least for 4.5 - 5.5. Almost all the crazy capabilities newer models have is from RLVR variants, whereas capabilities driven by RLHF are inching along.

    • This doesn't seem true.

      I'm not aware of any benchmarks that measure the first "analyze XYZ geopolitical situation" but legal reasoning is very closely related to "explain the ramifications of XYZ law".

      Legal Bench[1] measures legal reasoning. It's close to saturated (ie, there isn't a lot of room for improvement) but Fable scores 88% vs eg Opus 4.7 at 85%.

      There probably isn't a lot of room for improvement on something like this - there is enough disagreement in legal reasoning to mean 100% is going to be impossible.

      [1] https://www.vals.ai/benchmarks/legal_bench

  • How are the models making politics better? I don't count AI attack ads as an improvement.

    • Is this a serious question?

      Improvement in this context means "better quality results".

      You can use better quality models to do worse things with.

      I'm not making any claim about second order effects like that.

      1 reply →