Comment by ololobus
10 hours ago
I was looking at $10k Mac Studio with M5 Ultra and 256 GB for local experiments, but then struggled to find what really good modern model I can fit into it. Yes, it can run a good dense 27B at Q8 with plenty of context, but what beyond that? IIUC, some Deepseek flash variants at Q4 are also feasible, but I am not sure if the quality will be good. They also don’t run that fast, like about 30 t/s
So if I stay within 35B, especially MOE, my M5 Pro 64GB MBP can also run them well, and it can do plenty of other stuff too including gaming. While 256 GB with such RAM bandwidth and powerful GPU sounds like fun on paper, it doesn’t seem to be the next level compared to 64 GB
Really curious what people run on 256 GB Macs
I feel like for localAI t/s is less of an issue. Just make a PRD and run a ralph loop. For big slogging projects like reverse engineering, or converting a codebase to a new language it actually doesn't matter if it takes a day or seven days.
Yeah this is my experience. My 24GB 3090 + 64GB RAM takes a couple hours to crank out some code with largest Gemma 4 and Qwen3.8 models it can run
But in the meantime I get dishes done, vacuum, flip laundry... etc etc
Frontier models also seem in such a rush to emit anything they produce a mess that needs steering all day anyway
While I have not tested it, it feels like my local setup going slower is better at producing code that works the first time as its not trying to look fast for marketing sake
Sounds like it's worth waiting for M7 anyways, no point investing too much right now
https://news.ycombinator.com/item?id=48676795
At least wait until someone gets one in hand and post a review, and if it does work decently well there probably is going to be a long backlog.