Comment by weiran
6 hours ago
25% more usage sounds about right given the other token costs are down about 20%? I don't think cache read is a big portion of the overall cost.
6 hours ago
25% more usage sounds about right given the other token costs are down about 20%? I don't think cache read is a big portion of the overall cost.
Anything with long context quickly gets dominated by cache reads. Especially for interactive sessions I’ve got cache read % between 95% and 98%.
In my mix it's usually 98% or 99% at which point Fable 5.1 was pretty close to the same cost as Opus 5 due to the cheaper cached read. I've seen similar numbers for other people with long-running tasks running experiment loops and than sort of thing.
> I don't think cache read is a big portion of the overall cost.
For long running tasks it is. That's what made Deepseek so cheap.
Yeah, flash models, DeepSeek, MiMo, GLM, I love those things. For simple tasks like a daily routine shit, just setting up stuff and then doing the hard stuff in Claude/Codex, that's a reasonable approach for someone like me, a "gentleman code farmer", lol. And even lower tier stuff, I have the local models taking care of. Now that Jev is out I can finally have a true AI sysadmins managing my "cloud in the basement" homelab at the cost of electricity, which is not cheap btw