← Back to context

Comment by CuriouslyC

11 hours ago

My argument is that if you were to compare "analyze XYZ geopolitical situation" or "explain the ramifications of XYZ law" from Opus 3.5, 4.5 and 5.5, the difference would be marginal, at least for 4.5 - 5.5. Almost all the crazy capabilities newer models have is from RLVR variants, whereas capabilities driven by RLHF are inching along.

This doesn't seem true.

I'm not aware of any benchmarks that measure the first "analyze XYZ geopolitical situation" but legal reasoning is very closely related to "explain the ramifications of XYZ law".

Legal Bench[1] measures legal reasoning. It's close to saturated (ie, there isn't a lot of room for improvement) but Fable scores 88% vs eg Opus 4.7 at 85%.

There probably isn't a lot of room for improvement on something like this - there is enough disagreement in legal reasoning to mean 100% is going to be impossible.

[1] https://www.vals.ai/benchmarks/legal_bench