← Back to context

Comment by reilly3000

6 days ago

You can? It’s pretty remarkable what you can do with Qwen 3.6 on decent hardware, and Gemma 12b on nearly anything. If you can’t own it, take something for a spin on vast.ai where you can get all kinds of machines on tap for ~$0.20/hour+

llama.cpp has gotten so good that you can really get optimized performance with a few flags or none at all, and serve any model on-demand. Spend $45 to have your own GLM 5.2 machine for an hour. If you can’t manage full utilization you’d get 33.2M tokens, retailing for $128.

It takes some smarts to use dumb models, but renting time for frontier tokens is easy math. Vast also lets you run on spot which is great for training or agents with harnesses that can accommodate uncertain model availability.