Comment by Foobar8568
13 hours ago
Memory used : 38GB, and I haven't even started a LLM nor podman, I always fight with memory when using LLM on my mac with 48gb.
And I don't remember to have been able to have pushed to 200k context Qwen 3.6. 3.8 is running on my RTX 5090.
Qwen 27B runs very comfortably on a 5090. You need to use Q4 quants and Q8 KV cache. Here's the math
https://news.ycombinator.com/item?id=49514141