Comment by charcircuit

9 days ago

The model also does more of a brute force approach of trying a bunch of different things; testing theories, giving up, trying another thing. A lot of the LLM benchmarks do not care how long an LLM takes (as long it's below some upper bound in some cases). A human stopping to think can be faster compared to the model from going down a bunch of bad paths.