← Back to context

Comment by margalabargala

1 hour ago

> There is a lot of inefficiency right now keeping prices elevated. That will change very fast and soon paying for inference will likely resemble paying for hosting your app.

I don't know about "very fast" or "soon" unless you're speaking in geological terms.

SOTA models like Kimi 3 require thousands of GB of RAM/VRAM to run at speeds that are real-time useful. Manufacturing the memory necessary for that quantity to be available at app-hosting prices will take decades. Software efficiency solutions might drop needed memory by an order of magnitude in that time...but a tenth of an enormous amount is still pretty darn big so won't get us there "soon".