Comment by Imnimo
4 days ago
I dunno, "with tools" means different things for different models. It depends on what tools you give it access to. HLE demands a lot of specialized stuff. Like an interpreter for the esoteric programming language Piet for two questions. If you're not standardizing the set of tools, these aren't apples-to-apples numbers.
Even without tools it also outperforms Gemini 2.5 pro and o3, 25.4% compared to 21.6% and 21.0%. Although I wonder if any of the exam was leaked into the training set or if it was specifically trained to be good at benchmarks, llama 4 style.