Comment by tristanj

15 hours ago

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

Performance is significantly higher than Fable 5.1

Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)

> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.

The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.

With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.

  • Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)