Comment by spwa4
13 hours ago
Strange that the system card carefully seems to avoid any benchmark where you can also find scores for GLM, Qwen. There's barely any overlap with GPT 5.6 benchmarks. Just these:
Model HLE w/tools GDPval-AA v2
Claude Fable 5.1 65.0 1853
GPT-5.6 Sol 64.5 ~1711-1730
GLM-5.3 62.5 1769
DeepSeek V4 Pro 60.0 1590
Kimi K3 59.8 1682
Qwen3.8-Max 56.2 1739
No comments yet
Contribute on Hacker News ↗