Comment by villgax

17 hours ago

Lol, such a lazily written article by wafer.ai

GPUs. 8× MI355X (TP8) B300 (TP8+DCP8)

Decode tok/s per stream 118 tok/s 172 tok/s

Peak aggregate. 952 tok/s 1,568 tok/s

Peak aggregate per GPU 119 tok/s 196 tok/s

On every row the B300 beat the MI355X

The B200 is being forcefully compared against something which is not gonna fit within it's memory in a single node & not much details about multi-node interconnectivity, disagg or not. As expected of a shoddy slop.

The only point it won is of cost per hour is one aggregation website for rentals, the premium a B300 commands against the $3/hr AMD chip which no provider has in abundance. Never bothered to do TCO of owning the hardware either.

Did you see this section?

    To the B200’s defence, its numbers are somewhat deflated by the fact that it pays a cross-node all-reduce on the decode critical path (RoCE v2 at ~195 Gb/s) — it’s the only config here that spans two nodes