← Back to context

Comment by ManuelKiessling

5 hours ago

Thanks for the data!

Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.

Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?

So at FP16, I alone keep a 1,664 GiB system occupied all the time?

No, a cluster can server multiple users at the same time, providers cap the tok/s so that one cluster can run inference on multiple inputs at the same time. OpenAI with their new ultrafast mode is probably reserving the whole cluster or prioritizing requests of ultrafast users above others with a higher tok/s hence the high price and high speed. There's many other knobs providers tweak that they don't show the users, for example I doubt many providers are hosting the full FP16 version.

  • It's not based on rate limiting at all.

    The "expensive part" of generating the next token is streaming in the model weights from memory. The computations are relatively simple, which is called a "low arithmetic intensity" in industry jargon.

    So what they do is batch multiple chats together and compute the neuron activations for all of them together.

    This is vaguely similar to how some database engines work, where if multiple users need to run a "whole table scan" query, the additional users "join" the streaming workload of the first query mid-way, then loop back around to complete the first part that they missed. The AI accelerators don't do this looping, but the concept is the same: amortize the expensive I/O over multiple computations running in parallel.

    The "turbo mode" token rate thing is almost certainly your query getting sent to slower or faster hardware, like B200 vs newer B300 kit.

It depends hugely on what "rent usage of this model through one of the many LLM hosting providers" means. If you're asking them to host the model privately then yes, all of that 1.6T of RAM is likely in use holding weights, activations and KV cache by an inference engine that's only getting/answering requests from you alone. When you aren't actively using the model the hosting process is still active and waiting with all of that memory still wired to it.

As background: For the most part VRAM oversubscription/paging/swapping isn't a thing in the same way that RAM for a VM often is. There are some approaches to it, but (to my knowledge) not at that sort of scale.

There are some systemic reasons for this, but very broadly speaking the GPU vendors are building toward the highest bandwidth and lowest latency possible, and the overhead/complexity of something like protected memory modes serves neither of those priorities.