← Back to context

Comment by dudeinhawaii

4 hours ago

I have not experienced this (yet) but I have with models in the past.

I think it's important to have a solid benchmark where you KNOW there's a difference in model performance.

I have one around 3D modeling that models really land in the same space each time I run it. It's visual, and it's super clear. Sol has perpetually generated low quality work regardless of reasoning level. Astra was the first OpenAI model to suddenly leapfrog the pack and generate content that was production ready, beating out any other provider.

I haven't seen Astra regress (yet).

I think, if you want to be consistent and scientific about it, then you'd have to use the models via API and lock to a specific version. Via the subscriptions, you are floating on whatever the latest version is, vendor to vendor.