Comment by abtinf
6 hours ago
Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.
Per the link someone else posted, the actual difference in $/task is not nearly so stark:
https://artificialanalysis.ai/models/releases/claude-opus-5-...
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.
So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
A cheaper price has no value if I can’t use the thing I’m paying for.
The Claude lock-in simply disqualifies anthropic entirely (for my use).
> Even on a subscription you'll get considerably more usage out of Opus.
That's an incredibly bold assumption.
It's not, I have a subscription to both and Astra burns usage like crazy.
yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn't run out that quickly for me. it's great, but it's on the same tier as fable for me - use it sparingly, only when really necessary.
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
This is news to me. Excited to try it out! Thanks.
news to me as well. i thought you were forced to use Codex if you wanted their subscription. I completely ignored it because of that. How do we do it?
With the pi harness it just opens the browser (or gives you a link if you're on headless) and you sign on as usual.
ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
Yes. Also, you get image generation included with the ChatGPT subscription, which is very nice for certain kinds of development.
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
Can you give more details here? This sounds intriguing.
Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.
If you read the page, Opus is now significantly better than Astra while also being cheaper and having more performance headroom available.
I read the page. It seems like a marginal improvement.
Let's wait for independent benchmarks at least
the benchmarks provided are already from independent organizations:
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
https://artificialanalysis.ai/models/releases/claude-opus-5-...