Comment by cptskippy
6 hours ago
> Qwen-3.6-35B-A3B
The A3B models are super fast but I found the A3B Q4 model ran in circles a lot and ended up taking longer to complete tasks that 27B Q6 because it kept having to redo/rethink/fix something.
I was writing extensive prompts to rein it in and it would still ignore basic directives like "never force push on the repo, ask me instead". I ended up switching back to 27B after about a week of frustration and lost productivity.
We’re using 6bit quants since we have 32GB cards.
Gemma QAT is an honourable mention.