Comment by varispeed

4 days ago

If Fable gets correct answer quicker, then you might pay less than doing back and forth with Opus, plus you lose more of your own time.

I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish

same as it ever was. It seems your argument implies a belief that you should always use the best model. Others think that not all tasks require the absolute most powerful, expensive, model.

If doing a lot of heavy lifting there. Not only is it not a given that they'll get the correct answer for a lot of simpler tasks in fewer tokens, but smaller models are often available at far higher tokens/second inference.

There are certainly tasks where fable will be faster and/or cheaper, but there are plenty of tasks where even Haiku is as fast or faster and cheaper, or where you can e.g. get away with models like gpt-oss that you can get from inference providers providing 10x+ the token/second speed.

If you don't use enough tokens that relying only on Fable becomes a problem, then keep using just Fable. Personally, for my $200/week Max subscription I'd run out of the weekly quota for Fable in a day. At API pricing I'd go bankrupt if I tried doing the things I do with cheaper models using Fable.

The CursorBench plot, for example, shows that fable does have slightly better performance, but Opus is pretty close, and is less expensive per task