Comment by vlowther
6 hours ago
It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
I’m already able to use 1M context windows with the same machine than you and same model. Strange
Yeah, most of what I did was to add fused TQ support to leave more memory free for other nefarious purposes.