← Back to context

Comment by torginus

10 hours ago

You can see the breakdown here on what subtasks it outperforms and underperforms Fable.

For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark.

https://artificialanalysis.ai/models/gpt-6-astra

Edit:

Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too.