Comment by NitpickLawyer
16 hours ago
Check out the "Completed steps on..." graph in this [1] evaluation.
That graph gives a good perspective of what models they've tested, and roughly what "subject" each step covers. It is on that task that they note this:
> Kimi K3 reaches step 17 on average, compared with step 11 for GLM-5.2.
[1] - https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...
No comments yet
Contribute on Hacker News ↗