Comment by akulbe

14 days ago

How well do you folks think this would run on this Apple Silicon setup?

MacBook Pro M2 Max

96GB of RAM

and which model should I try (if at all)?

The alternative is a VM w/dual 3090s set up with PCI passthrough.

Depends on quantization. 109B at 4-bit quantization would be ~55GB of ram for parameters in theory, plus overhead of the KV cache which for even modest context windows could jump total to 90GB or something.

Curious to here other input here. A bit out of touch with recent advancements in context window / KV cache ram usage