← Back to context

Comment by londons_explore

8 hours ago

Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)

Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.

> ~1000 bytes per context token per user

Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb