Comment by user43928
1 hour ago
It's an interesting question. Some napkin math:
Let's say we serve a Fable class model on 8x B300.
From Kimi K3 metrics, with 8x concurrent streams, we would achieve 55-60 tok/s per stream, matching Fable 5.1 throughput.
432 tok/s x 3600 => 1.555M output tokens/h x 50$/M API price = $77.76 revenue per hour.
Assuming total API billing at 2.06x output token bill = $160.2 / hour or ~$20 per B300.
A server with 8x B300 could be $461.5k.
At an obviously unrealistic 100% utilization we would look at 4 months of revenue to match the cost of the server.
About how model serving works at scale and actual utilization I know little.
And for all we know Anthropic could serve their model with 64 streams on the same hardware instead of the 8 we assumed here.
No comments yet
Contribute on Hacker News ↗