Comment by jodleif 2 days ago Then you might be missing SWA. Gemma models are extremely memory hungry without 2 comments jodleif Reply CMay 2 days ago So long as they have flash attention enabled, Llama.cpp enables Sliding Window Attention by default for Gemma 4 models. Even if they're using Ollama or LM Studio I would expect those to mostly be doing the right things. DiabloD3 2 days ago I would not expect Ollama to be doing the right thing fwiw.
CMay 2 days ago So long as they have flash attention enabled, Llama.cpp enables Sliding Window Attention by default for Gemma 4 models. Even if they're using Ollama or LM Studio I would expect those to mostly be doing the right things. DiabloD3 2 days ago I would not expect Ollama to be doing the right thing fwiw.
So long as they have flash attention enabled, Llama.cpp enables Sliding Window Attention by default for Gemma 4 models. Even if they're using Ollama or LM Studio I would expect those to mostly be doing the right things.
I would not expect Ollama to be doing the right thing fwiw.