Comment by Eridrus

5 hours ago

It's actually existing flat per-token pricing that is weird.

Neither encode nor decode are linear in compute, so providers need to price for average expected length.

This is just getting closer to the true cost of generating tokens.

My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.

Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.

Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO