Comment by jasonjmcghee
14 days ago
Depends on quantization. 109B at 4-bit quantization would be ~55GB of ram for parameters in theory, plus overhead of the KV cache which for even modest context windows could jump total to 90GB or something.
Curious to here other input here. A bit out of touch with recent advancements in context window / KV cache ram usage
No comments yet
Contribute on Hacker News ↗