Comment by tintor
5 hours ago
- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/
- Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra
Who is wrong here?
Some benchmark results in Astra page for Fable and Opus are blank (-).
What is Artificial Analysis intelligence index measuring that Astra scores poorly on?
Can someone from OpenAI / Artificial Analysis comment / clarify?
Even OpenAI Astra page mentions the low scope from Artificial Analysis for Astra.
I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is.
That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.
But AA scores Gemini 3.8 Flash at 59, and Astra at 61.
Many people claim that the Artificial Analysis Index is highly contaminated - I have not personally looked into it.
Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.
OpenAI clearly cares about Artificial Analysis Index since they included Astra score from Artificial Analysis Index.
> Humanity's Last Exam (w/ tools)
This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended research-level questions. It tests domain knowledge and problem solving skills. Burning more reasoning tokens may help somewhat but not as much as e.g. coding benchmarks.
If you scroll down in the Artificial Analysis page you linked, you'll see all the individual benchmarks.