Comment by edg5000
6 hours ago
Cache reads dominate in modern workflows (coding CLIs and modern web clients such as ChatGPT Work and Claude Cowork (web)).
6 hours ago
Cache reads dominate in modern workflows (coding CLIs and modern web clients such as ChatGPT Work and Claude Cowork (web)).
Output tokens are 5x more expensive than input tokens, so I'm not sure "dominate" is entirely correct.
A conversation with 20 turns, 50k tok growth per turn, 1m tok context at end would price out like this:
Fable 5 ($1/M cache reads) ; cache reads 9.5M tok × $1.00 = $9.50 ; cache writes 1M tok × $12.50 = $12.50 ; output 1M tok × $50 = $50.00 ; total = $72.00
Fable 5.1 ($0.25/M cache reads) ; cache reads 9.5M tok × $0.25 = $2.38 ; cache writes 1M tok × $12.50 = $12.50 ; output 1M tok × $50 = $50.00 ; total = $64.88
So yes, cheaper, but not massively.
[dead]