← Back to context

Comment by Roark66

1 hour ago

Bingo. They are scared s*itless of local models. Soon the raising interest rates will mean they can't just buy the entire world's gpu/ram/storage capacity and lock it away in a dark room for no one else to use them.

Once prices of hardware stop being bonkers many people will run local.

People often do those silly counts looking at local generation speeds of 100tok/s and saying you can only get 8M tokens a day and that is worth so little you'll never offset the local hardware cost.

But they are forgetting about two huge things. One, the split between input/output token use is huge in typical programming. I typically use 1.2-1.6B (as in Billion) input tokens and only 8M output a week. Out of that 75-80% of input is cached on Anthropic with their pretty inflexible short lived cache.

And here local AI shines. You can save your contexts to disk so you can go back to a session 3 weeks later and load it all from cache without having to prefill. If you have the RAM and you use models like Qwen4 (3.8 flash next) that can fit 6 to 14 262k contexts in 48GB of system ram as cache.

And the second thing is you can have 6 to 14 sessions that do 99% caching and you can run a lot more input for a long time in 48gb dedicated system ram. (the number varies a bit depending on the content of the context).

So you have RAM that caches "automatically" and you can share the cache between users. Or if you can remember to save to disk (server side, the client just sends a request to save/load). But this uses both RAM and flash storage. Two things that are horribly expensive now.

However when you have a local model it enables workloads that were simply completely impossible in the cloud.