Comment by Iolaum
7 hours ago
I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
7 hours ago
I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
https://openai.com/index/our-decision-on-cursor-following-it...
Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.
They have Astra in other benchmarks lower on the page. They just don't want to show it winning
The chart is cursorbench though and they asked about the "deceptive graph"
They can benchmark it because you can use an openai api key with cursor. Astra is just not included in the cursor plan.
Elon posted on X that Grok 4.7 is behind Claude and OpenAI for agentic coding:
https://x.com/elonmusk/status/2102082011233931762?s=20
so it's likely about usage in Cursor specifically.
https://xxcancel.com/elonmusk/status/2102082011233931762?s=2...
this is exactly it.