← Back to context

Comment by nl

4 hours ago

This doesn't seem true.

I'm not aware of any benchmarks that measure the first "analyze XYZ geopolitical situation" but legal reasoning is very closely related to "explain the ramifications of XYZ law".

Legal Bench[1] measures legal reasoning. It's close to saturated (ie, there isn't a lot of room for improvement) but Fable scores 88% vs eg Opus 4.7 at 85%.

There probably isn't a lot of room for improvement on something like this - there is enough disagreement in legal reasoning to mean 100% is going to be impossible.

[1] https://www.vals.ai/benchmarks/legal_bench

That benchmark agrees with my point, GPT-4o scored an 80% on it. Even if you could construct a much harder legal bench that could resolve smarter models, the bottleneck is getting the absolute top legal experts in the world to give feedback, so progress would be slow compared to stuff like math and coding where you can programmatically generate scenarios and validate/score performance.

These models are spikey as hell, they can gain incredible capabilities but that doesn't make them godlike minds, it's more like a scaled up version of rain man.