Comment by segmondy
1 day ago
I wonder if some of these could carry over for LLM inference. It will be nice to turn more ewaste GPUs into capable processing units for LLM.
1 day ago
I wonder if some of these could carry over for LLM inference. It will be nice to turn more ewaste GPUs into capable processing units for LLM.
I have a Radeon RX 6900 XT (16 GB VRAM, originally released in 2021), and it's possible to run some lightweight models to have okayish performance and quality of output, but nothing I've tried has come anywhere close to the quality even of the models I can use for free from OpenCode Zen or the free tier of Openrouter. If you want to keep everything local on the same card I have, it requires putting up with a model that's noticeably worse in virtually every metric than what you can get for free elsewhere, and the GPUs this article are talking about are three times as old as mine.
It would be awesome if someone manages to figure out how to get small enough models to fit on older cards to be viable, but I'm not optimistic that it will come without some sort of fundamental architectural innovation rather than incremental improvements, and it's not clear if and when that will happen.
This kinda highlights the level of debt the AI companies are in, and will continue to be in, offering anything for free.
How long is this runway?
I have no clue, and I agree that it does not seem sustainable. Either someone needs to find a magic solution to making it a lot cheaper, or a lot of companies are going to need a lot of money from somewhere that isn't clear.
Define "small".
The other day I managed to get a context of 195k for Qwen3.5-9b Q4_K_M using a llama.cpp fork that supports TurboQuant:
https://github.com/TheTom/llama-cpp-turboquant
I think you could replicate this with a larger model on your device.
Overall with the right quantisations for both the model and KV cache you can get a lot of mileage out of this old hardware. Speed remains the main limitation, as IIRC I was getting ~26-30tps on a 7700S.
Opencode currently has Mimo 2.6 Flash Free with 5x that context. Like I said, it's not impossible to get something usable, it's just not in the same league as what you can get without paying anything or needing to spend effort "managing" to get it working.
I'm not saying there's no benefit to using local models. My point is still the same as before: you have to be willing to sacrifice both performance and quality even when just comparing to free models that are available today.
1 reply →
With an extra 8gb of vram you could run qwen 3.8 27b pretty comfortably, which isn't quite as good as frontier models but definitely on par with free models on openrouter and whatnot.
Also you can use multiple GPUs at once, two of your GPUs could run qwen 27b very comfortably, and with great performance.
OpenRouter has Qwen 3.8 27B in its free tier right now. My point was that if a 5 year old GPU can't even compete with the free offerings, the 15 year old ones mentioned in the article aren't going to come close.
I have a 6900 XT as well. Unfortunately it is only 512 GB/s.
With AMD the best you can do in the consumer market right now is an RX 7900 XTX which is about 960 GB/s.
Most of these older cards are lacking the physical hardware for fp8 or other lower precisions that most quantized models use. Or the memory to run models at higher precision.
Memory is the bottleneck along with the lack of FP4/FP8 capability at the hardware level.