← Back to context

Comment by user43928

2 hours ago

It's an interesting question. Some napkin math:

Let's say we serve a Fable class model on 8x B300.

From Kimi K3 metrics, with 8x concurrent streams, we would achieve 55-60 tok/s per stream, matching Fable 5.1 throughput.

432 tok/s x 3600 => 1.555M output tokens/h x 50$/M API price = $77.76 revenue per hour.

Assuming total API billing at 2.06x output token bill = $160.2 / hour or ~$20 per B300.

A server with 8x B300 could be $461.5k.

At an obviously unrealistic 100% utilization we would look at 4 months of revenue to match the cost of the server.

About how model serving works at scale and actual utilization I know little.

And for all we know Anthropic could serve their model with 64 streams on the same hardware instead of the 8 we assumed here.