Comment by jasonjmcghee

14 days ago

Depends on quantization. 109B at 4-bit quantization would be ~55GB of ram for parameters in theory, plus overhead of the KV cache which for even modest context windows could jump total to 90GB or something.

Curious to here other input here. A bit out of touch with recent advancements in context window / KV cache ram usage