Comment by jodleif

2 days ago

Then you might be missing SWA. Gemma models are extremely memory hungry without

So long as they have flash attention enabled, Llama.cpp enables Sliding Window Attention by default for Gemma 4 models. Even if they're using Ollama or LM Studio I would expect those to mostly be doing the right things.