← Back to context

Comment by abletonlive

2 hours ago

> allow me to educate you

No thanks, you're not in a position to do that clearly.

> Since you clearly don't use local llms

I do, probably a lot longer than you have actually.

> anything under 100 tok/sec is USELESS

Objectively wrong. You sound like you're really behind and you're so myopic that you think coding is the only use case for local LLMs. I'm a professional software dev and that's the least interesting use case of local LLMs.

> Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec.

You clearly didn't read the article or have reading comprehension issues. The model is Qwen3.8-Flash-Next 4 and 5-bit quant, neither of which "conveniently fits on one GPU". Sorry that your hardware doesn't live up to your own delusions and can't even run Qwen3.8-Flash-Next at 4/5 bit quant. You are taking the Quen3.8-27B numbers, something that the article isn't really that concerned with, and trying to make it fit into your narrative.

> So I ask you again, which one are you?

Well I'm someone that suggests that you should touch some grass and reevaluate your personal issues. You seem angry. Perhaps it's best to figure your own issues before trying to figure out why people are excited about Apple hardware for local llms. I am sure the people that need to interact with you in society would be very grateful if you took the time to do this.