Comment by msabalau
4 days ago
Not everyone is you. Other people probably have a range of tasks that can accomplished with different models.
Certainly if I'm confident that I'm going to get what I need from a faster model, that's what I want to use, rather than wasting time grinding away for the sake of saying of the same answer came from a SOTA model.
Given that every chatbot does offer a range of models, it seems clear people do choose among options.
The mental effort in estimating what model would be better is so not worth it.
I just want to switch to Claude Code, tell it to turn a .csv into a BigQuery table then cmd+tab to something else while it runs. Thinking "oh this is probably an easy task, I can /model to Sonnet to save $0.0004" is silly.
People actually use the models for more than writing code.
I barely used Fable because of the rate limits. It just makes more sense to use Opus.
If there were no limits it would be different.
It is not to save a fraction of a penny, it is to be able to still use the model within the limit on the week for $20.
Not everything is interactive, and what the routers do is precisely take away that mental effort on repeated tasks.
The game changes when it's not just "Opus vs Sonnet", but "Opus vs GLM". The amount saved is way more than even $0.04. And it's not only money but speed. Some providers can serve GLM crazy fast - I'll even go outside of my subscription to pay extra money for the quick results.