Comment by 7speter

6 months ago

Elsewhere in the thread, someone talked about how h100’s each have 80GB of vram and cost 20000 dollars.

The largest chatgpt models are maybe 1-1.5tb in size and all of that needs to load into pooled vram. That sounds daunting, but a company like open ai has countless machines that have enough of these datacenter grade gpus with gobs of vram pooled together to run their big models.

Inference is also pretty cheap, especially when a model can comfortably fit in a pool of vram. Its not that the pool of gpus spool up each time someone sends a request, but whats more likely is that there’s a queue to f requests from someone like chatgpts 700 million users, and the multiple (I have no idea how many) pools of vram keep the models in their memory to chew through that nearly perpetual queue of requests.

0 comments

7speter

No comments yet

Contribute on Hacker News ↗