Comment by andy_ppp

2 months ago

Honestly you're still looking at (from my understanding) ~3 minutes prefill (TTFT) even with architectural improvements and so on with a 32k context window (against a large model). How is this going to be competitive with Nvidia and all of the tricks massive scale get's you to parallelise context across many machines?

Is it supposed to be? I think the point with some of these Macs is you get the capability in something the size of a heatsink from Intel's Netburst architecture era, or a Macbook light enough to stick in a backpack and take with you to lunch.

If you're talking about chaining together multiple GPUs you're talking about a different game -- I suspect, anyway. Seems like a high-spec Mac would be good for development and testing. Arrays of GPUs, better aimed at production use.