Comment by Tade0
12 hours ago
Define "small".
The other day I managed to get a context of 195k for Qwen3.5-9b Q4_K_M using a llama.cpp fork that supports TurboQuant:
https://github.com/TheTom/llama-cpp-turboquant
I think you could replicate this with a larger model on your device.
Overall with the right quantisations for both the model and KV cache you can get a lot of mileage out of this old hardware. Speed remains the main limitation, as IIRC I was getting ~26-30tps on a 7700S.
Opencode currently has Mimo 2.6 Flash Free with 5x that context. Like I said, it's not impossible to get something usable, it's just not in the same league as what you can get without paying anything or needing to spend effort "managing" to get it working.
I'm not saying there's no benefit to using local models. My point is still the same as before: you have to be willing to sacrifice both performance and quality even when just comparing to free models that are available today.
Yeah, but some prefer this if it means they own the infrastructure.
I spent some years in the pharma industry, where AI came really late, as there was (justified) concern that sensitive data might leak - even by accident. Eventually a solution was implemented - it was some kind of open model (they didn't say which) running on-premises.
You couldn't tell these people that this and that powerful model is free or inexpensive, because it's useless to them if it's running on someone else's computer i.e. the cloud.