← Back to context

Comment by oblio

2 days ago

> and make inference cheap enough to eventually escape the red numbers

Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ?

We do know about the hardware needed for a given token speed. What that hardware costs, and electricity prices.

With that its easy calculations to get about the profit margins for a given price for a given model.

  • I'll plug your comment into a couple LLMs. If it's so easy, they should be able to provide the numbers.

    Edit: Gemini 3.5 Pro and Opus Claude 4.8 both disagreed that it's easy to determine anything, for both the high end (Fable, Sol) or the low end (open weights). Due to competition, subsidies, etc, gross margins could be as high as 85% (extremely unlikely) to as low as 10% or even negative. And that's just for pure inference and gross margins. Even for pure inference providers this doesn't include any overhead such as rent for office space for the pesky humans operating the business, marketing expenses, etc, etc. Let alone any crazy soul that actually wants or needs to train something.

    • A lot of datacenter are operating their own Natural gas based power generation. Which in itself is a different dynamic than buying electricity.

      I have operated such a setup, at a much lower industrial manufacturing scale. The tradeoffs are quite stark. The electricity is cheaper, but generators/turbines need to operate at 80% capacity to be feasible. In the slow hours, they become an albatross.

      So they lose more money per user if less people are using the services, but they also lose money overall if more people are using them.

Why would bedrock sell at a loss?

  • Because Amazon needs to justify $175bn of yearly capex spending and $2.5tn of market cap? Amazon owns a big chunk of Anthropic and a bit of OpenAI.

    For Magnificent 7 the AI bubble bursting will probably wipe out 30-50% of their valuations until the next tech cycle begins.