Comment by TechSquidTV
12 hours ago
In my limited experience, not quite yet but we are damn close. Qwen 3.8 27b is it. If I could run this as a decent speed, I would no longer need cloud models at all. I'm actually currently trying it out in the cloud to pay for the inference speed but the model is fully runnable at home.
I realistically costs $5-10k to replicate a ChatGPT like agent. And it doesn't scale.
That's still really close. And models and quantization etc keep improving.
I'm absolutely positive that I'll be switching to mostly local AI in the next 5 years.
I could believe that especially with some of the chips playing catch-up.
E.g., the M7 chip is rumored to be the one where Apple has poured really serious effort into local AI performance where previous generations seem to have mostly been coincidentally good at it.
Maybe this is an incorrect opinion but I don’t personally think that the M1-M3 or maybe M4 or even M5 chips were designed with LLM inference in mind at all. These were designed with things like video rendering, image/video ML, and rasterization performance in mind.
Good 60-70 tg and 2K pp Qwen 3.8 27B FP8 can be had for about 5-6K (2xR9700 + PC) Gives about 3-4 concurrent sessions with full 262K
Fast 150+ tg and 2-8K pp Qwen 3.8 27B nvfp4 is about 8K (5090 +PC) Gives really only one concurrent session that flies because kv caching is not perfect for ninfer https://github.com/Neroued/ninfer
both are very serviceable, I prefer FP8 on 2xR9700
But, yes it doesn't scale that well but in 5 years the same hardware should still be very capable of running some great MoE models, for example Qwen 3.6 35BA3B on 5090 can fly at 600 tg
Qwen 3.8 at 27b, 4bit MTP, Full context, in 72GB blackwell is 2-3x agents.
If you have a real product and can actually sell it, youre taking a largish risk relying on the cloud.
From model changes, alignment, to enshittification and the natural cognitive offloading, you could be one day removed and ROI tanked.
Think of AI like a mafia boss who helpfully supports you untill they need a favor. Thats all cloud AI is in America.