Comment by try-working
4 hours ago
There are insane speed improvements for local inference going around on X right now. They've popped up the last month and week.
Tensorfold is getting 100%+ speed increases on both prefill and decode for models like Qwen 27B. oMLX has followed them and have had similar improvements in the past week.
There's lots of different techniques like letting CPU help with prefill, DFlash specualtive decoding etc.
I'm really excited for this as I'll be receiving an M5U in about a month. Expect to be running Qwen 4 27B or Flash (it's a 96gb machine), and they may come close in performance to DS 4/4.1 Flash, and should be able to hit 100 tps. Local is really becoming viable, especially considering that GPT 6.1 has been running at 20ish tps the past week.
No comments yet
Contribute on Hacker News ↗