Comment by visarga

18 hours ago

Large LLMs on MacBook produce tokens at an acceptable speed but the problem is reading context. Not incremental reading like when you have a chat session, because they use KV cache, but large size reading, like when you paste a big file. It can take minutes.

7 comments

visarga

antirez 17 hours ago

DS4 can process 460 prompt tokens per second. Not stellar but not so slow. On M3 max. See the benchmarks on readme.

habosa 9 hours ago

Can you ELI5 why this is so slow for local inference but so fast for using hosted models?

bel8 18 hours ago

And unless I'm mistaken, the repo is about running it with 2bit quantization.

This is probably far from the raw intelligence provided by cloud providers.

Still, this shines more light on local LLMs for agentic workflows.

antirez 16 hours ago
It runs both q2 and original (4 bit routed experts). At the same speed more or less. The q2 quants are not what you could expect: it works extremely well for a few reasons. For the full model you need a Mac with 256GB.
- someone13 10 hours ago
  
  Out of curiosity, do you have any theories of why it works so well at such aggressive quantization levels?
  
  1 reply →

brcmthrowaway 17 hours ago

Why is this the case?

Are there any architectures that don't rely on feeding the entire history back into the chat?

Recurrent LLMs?