Comment by bigyabai

10 hours ago

It's likely that a stack of Macs will draw more power for slower prefill/decode than equivalently priced Nvidia GPUs. If power efficient inference is the goal, Macs are a non-starter.

So if it isn't a comparative ability, now it's a power cost issue? This reads like goal post moving.

  • Oh, it's absolutely both. The power you waste waiting for TFTT on prefill will absolutely compound at the "medium sized labs and businesses" scale.