Comment by tobiasu
2 hours ago
Sure. The fastest small coding model is probably Mellum2 12B-A2.5 by Jetbrains. It matches or beats all Qwen models in this class.
Can even run on a notebook CPU and comes in Base (best for FIM), Instruct and Thinking variants. mradermacher has imatrix quants for people who can't run it at Q8.
IQ4 should fit, but even if it doesn't, llama.cpp has options to partially offload models to system memory.
No comments yet
Contribute on Hacker News ↗