Comment by kenmu
5 days ago
Is their scaffold available? Does it do anything special beyond feeding the warmup, single challenge, and full problem to an LLM? Because it's interesting that GPT-5.2 Pro, arguably the best model until a few months ago, couldn't even solve the warmup. And now every frontier model can solve the full problem. Even the non-Pro GPT-5.4. Also strange that Gemini 3 Deep Think couldn't solve it, whereas Gemini 3.1 Pro could. I read that Deep Think is based on 3.1 Pro. Is that correct?
I see that GPT-5.2 Pro and Gemini 3 Deep Think simply had the problems entered into the prompt. Whereas the rest of the models had a decent amount of context, tips, and ideas prefaced to the problem. Were the newer models not able to solve this problem without that help?
Anyway, impressive result regardless of whether previous models could've also solved it and whether the extra context was necessary.
I know these frontier models behave differently from each other. I wonder how many problems they could solve combining efforts.
No comments yet
Contribute on Hacker News ↗