Comment by CuriouslyC
4 hours ago
That benchmark agrees with my point, GPT-4o scored an 80% on it. Even if you could construct a much harder legal bench that could resolve smarter models, the bottleneck is getting the absolute top legal experts in the world to give feedback, so progress would be slow compared to stuff like math and coding where you can programmatically generate scenarios and validate/score performance.
These models are spikey as hell, they can gain incredible capabilities but that doesn't make them godlike minds, it's more like a scaled up version of rain man.
Agree about the spikeiness, not sure about the legal example.
I think there is a roof on the level of reasoning needed for legal reasoning, and I think we are pretty close to it. Legal arguments just aren't that complex in terms of logic chains.
I'd bet the cyclometric complexity of any given court case is lower than a dense piece of code for example.