← Back to context

Comment by pipsterwo

8 hours ago

1/1000 of inference compute is a non-trivial workload at scale. Gartner estimates ~$28B in inference spend for 2026 making this a $28 million dollar per year workload (edit: based on the assumption above)

Source: https://www.gartner.com/en/newsroom/press-releases/2026-07-2...

The issue is it’s cpu compute which is underutilized in gpu clusters anyway, so practically it’s not really 1/1000.

  • Totally, edited my comment to specify "based on the assumption above." The main takeaway I was going for was 0.1% is not a small number in this context