← Back to context

Comment by mediaman

7 days ago

The (quite excellent) article discusses several of your points. If you haven't read it, I recommend it.

  - Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.

  - Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use

  - The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6.

  - This is because US labs are leading on cost efficacy of inference ($/task)

  - Training will decline as a percentage of costs as inference expands compute share due to agentic workloads. A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.

  - With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates

OpenAI really shows the way here. Their cost per task is less than half that of Anthropic because of more efficient tokenization and less verbosity. OpenAI is both cheaper and better than Chinese models for frontier work.

Ben's article "distills" down to 2 reasons that US frontier labs shouldn't be "afraid":

1. US frontier lab unit economics are better 2. US frontier labs are moving up the stack making tools that are "stickiness" and will prevent users from switching.

For 1...he doesn't provide any evidence for US lab unit economics being better...the major input to unit economics is electricity...which is cheaper in China. And building data centers and connecting them to electricity is both cheaper and an order of magnitude faster in China. The main input that US labs might have an advantage in is in cost/access to chips, but that given the level of chip investment in China it seems unlikely to hold.

For 2...there's little evidence these tools are sticky. At least in programming, the trend seems to be tools like opencode that support multiple models and providers.

And even when they are sort of sticky, as we know on hacker news, people figure out how to point the tools they like to competing models even when the app doesn't official support it.

And every improvement in model capability makes it increasingly easier to make your own tools.

Wrote more on this in a blog post that has an earlier HN discussion: https://larrysalibra.com/ben-thompson-is-wrong-us-frontier-l...

  • He’s glossing over the reason they are not: 90% profit margin of Nvidia. Power is only a small part, single digit, it will eventually matter but does not really today.

    What is the cost of AI? The single largest ingredient is Nvidia profit margin.

    Huawei accelerators are not as efficiency yet, but they don’t nearly extract as much margin.

    Why would future revenue stay with the labs given this situation? This whole thing had an airline industry sized red flag on it that makes investing into frontier lab about as sexy as investing in United.

    Maybe the token economy is some kind of reverberation of the airline reward miles economy, the emergency hatch to be able to survive under maximal supplier extraction (Nvidia is just the top of a monopoly stack here, even if they replace those chips, the HBM, ASML, Foundry layer can get their dues)

    • > Power is only a small part, single digit, it will eventually matter but does not really today.

      Sorta yes, sorta no.

      A single 5090 consumes 450W - at Californian energy prices of $0.38 per kWh that's $0.17 per hour. And the card itself costs $4100 on amazon. So after 2.75 years running at full power 24/7 you'll have spent more on electricity than on the card. I would have thought most data centres being built today would have a design life longer than 3 years.

      Of course you can throttle the cards to ~300W without losing too much performance. But also you need more than a single 24GB card to run most modern LLMs.

      5 replies →

  • Cost of electricity isn’t a long term advantage in my opinion. Private companies will figure it out.

    What matters most is $/completed task. It does seem like OpenAI and Anthropic are winning here even with worse electricity rates. Perhaps it is made up by the efficiency of Nvidia and Broadcom chips, which China can’t get in mass.

    I do think that OpenAI and Anthropic are moving up in stickiness. My company has rallied around Claude. We are customizing Claude Code, adding knowledge bases for non technical people, writing skills for them, using Claude features company wide. It’s hard to move.

    Meanwhile, I personally use ChatGPT outside of work. The memory, ease of use, habit keeps my subscribed.

    • I'm a model nomad, using whatever solved my last problem the best and where it makes the most sense to start my next work in.

      However with the latest models Fable, Kimi K3, 5.6, it's getting to a point where I sometimes forget what model I am on without noticing a difference. And once I realize it because something may not be exactly like I expected it I won't switch for that work either because I don't want to invalidate the cache.

      For the next work I will do there is maybe a 50/50 chance to remember to switch the model before I start.

      That's not what I would call stickiness towards a certain provider.

      3 replies →

    • Seems like most popular harnesses, including codex and Claude code, support Agent Skills (an open spec for skill formatting/ organization): https://agentskills.io/clients

      Which is to say, this isn't really a lock-in/ stickiness vector (unless maybe the wording itself of a skill is hyper-optimized for a specific model)

      1 reply →

    • > Private companies will figure it out.

      Across sectors, China added 543 GW of energy in 2025. Next year, USA is expected to add between 70 and 80 GW of energy

      6 replies →

  • > 1. US frontier lab unit economics are better

    That's not generally true, since there is generally still much reliance on NVIDIA. The true low cost providers are Google with their TPU and vertically optimized stack, and Amazon with Trainium. However, Google does not have their own frontier model, and Anthropic (who are partially served by Amazon) are also paying a premium for extra NVIDIA-based capacity from SpaceX, maybe soon from Meta too.

    I don't know how the economics of domestic Chinese Huawei-based clouds (no NVIDIA) compares to the west, but since serving cost is mostly hardware depreciation and to a lesser extent electricity, they are not necessarily at a disadvantage (Ascend 950 costs roughly 50% of an NVIDIA H100), and more to the point it is irrelevant when considering US commercial use that is more likely to be using Chinese open weights models from US providers served on NVIDIA based hardware.

    I think the real significance of Chinese frontier models being open weight is that it takes development cost amortization out of the US-based serving cost, while the US AI labs can't afford to do this. The US labs therefore need to reduce development spending to remain price competitive. The Chinese companies are of course still making money from the Chinese market, whether by selling API access or by other business models such as Ziphu making 75% of it's total revenue by selling services to Chinese customers who are running their models on-prem due to the Chinese apparently being very concerned about data privacy.

  • More basically, production cost matters only if inference is priced at commodity prices. That's not what VC's signed up for, which is rent-seeking.

  • In a corporate setting yes Opencode all the way. However in a non corporate setting I am getting $3000 of api usage a month for $100 at Anthropic and only use open code for the smallest cheapest tasks

Sure, let's have a look...

> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence. [emphasis mine]

I guess I'm missing the part of this article where they bring hard numbers in to back up the argument here. What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively. Composer 2.5 is itself a fine-tune of K2.7, and it's apparently quite token efficient, so why would it be impossible for a Chinese lab to achieve something similar? GLM 5.2 Max is also ranked above the lower end OpenAI models and is not far off in price.

It's weird to have this entire discussion about tokenomics without mention of the circular financing and debt raised by labs in the West, which can then essentially give away their capacity to end users. OpenAI giving away quota resets to subscribers like candy on Halloween while their compute partner Oracle's bonds is reevaluated to be one grade above junk? How?

I don't think you can make an argument about the future one way or another by arguing using the listed prices. The math is not internally consistent enough for it.

  • >What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively

    Because you're comparing retail price whereas the parent commenter (and the article) is talking about marginal (ie. inference) costs. American labs are providing a premium product and they're charging accordingly. Meanwhile for chinese models they're open weight so they're limited to how much they can charge without competitors undercutting them.

    If we use tokens as a rough proxy of inference costs (rough approximation, I know) and look at artifical analysis benchmarks, you see that all the open models are behind the pareto frontier in terms of efficiency.

    • I'm arguing we can't trust retail prices because the marginal pricing isn't meaningfully connected to it anyway.

      But if we have to look at what we think margins might look like, DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price (https://openrouter.ai/deepseek/deepseek-v4-flash), so there's at least one example of a Chinese lab charging a predetermined price despite competition. And no one but Moonshot is hosting Kimi K3 yet (https://openrouter.ai/moonshotai/kimi-k3). Perhaps there's room in the market for those who release their models to make margin on them.

      And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.

      6 replies →

US running costs are higher than in China, because the US lags behind in energy, has higher real estate costs, and wage costs are higher.

Eventually we will hit a "good enough for cheap enough" and frontier models will hit diminishing returns (if they haven't already for a lot of types of work)

Don't think the rest of the world will sit on their hands while the US soaks up chips either, demand gets filled and if the US won't fill global demand for chips that's an opportunity to undercut again.

The other thing the rest of the world doesn't have to fund is the ridiculous valuations on these companies.

Unless you think the US can stay ahead just with model efficiencies, and that no one else will eventually match them, you are looking at the writing on the wall.

All that to say, the rest of the world is more than willing to eat your lunch, they have a dozen good reasons to, and they're already showing good results.

Just on the economics side, we've been here before too, US companies typically export their commoditization and live on brand royalties. Think all the cheap manufactured goods, the US doesn't make any of it. That's because the US can't compete on margins for numerous reasons, it's too expensive, I don't think AI is any different here except that the brands are currently valued in the trillions and I suspect that greed will be their undoing.

  • The US does not lag behind in energy. Industrial electricity prices in most places in the US are competitive with China, or even cheaper.

    • Apologies, I meant ability to deploy new generation. You're right that areas of the US are cost competitive.

      This can change quickly though, so it's not that big of a deal. If AI energy demands push the Gov to deregulate/fast track new plants, or the industry decides to build out their own generation renewables.

I think there are some really interesting thought there, but I’d challenge some of this:

> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use

I think a large part of manufacturing economics is illiquid overhead and the cost of expertise to set up and run your manufacturing line. Compute economics don’t have the same illiquidity nor do they require the same expertise or even specialized infra (current temporary chip shortage aside).

The implications of this are small players (e.g. your uncle running an inference server out of his garage) have comparably efficient marginal costs as big players. Compare this to actual manufacturing where small players have essentially no access to the manufacturing facilities of the big players.

Additionally, big players with a lot of compute who are not meaningfully in inference today (e.g. Amazon) have a fairly straightforward glide path to utilizing that compute to compete.

> This is because US labs are leading on cost efficacy of inference ($/task)

It’s possible, but I would need to see better data on this.

>A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.

I think it’s fair to assume this is true, but also token efficiency is not a meaningful competitive moat. It’s not like these are secrets the Chinese will never figure out, it’s a fairly active research space and the outcomes are quantifiable.

  • Amazon is not "meaningfully in inference"? Bedrock seems to have a ton of enterprise customers, some of which would never trust the AI labs themselves with their data but will trust Amazon.

    • Relative to their other compute or other players inference, not as significantly. Though yes, they certainly have some share.

    • A lot of this companies are also in parallel evaluating running Chinese models inhouse to be independent of the good will and profit margins of externals.

      This is where Oracles "datacenters for rent to run your own Chinese models" strategy will benefit. The LLM SaaS game is a lost one thanks to China.

But this assumes Chinese models will not achieve token cost optimization. Intelligence needs are fairly flat for many tasks, and the Chinese models have caught up on this front. Next they achieve greater token cost efficiency and we don’t need OpenAI.

  • That the author doesn't acknowledge the relentless R&D efforts DeepSeek has been plowing into optimization, and giving a default win to OpenAI/Anthropic on the supposition that they've been serving models for longer is a black mark against the article.

    I appreciate the transparency in explicitly stating their motivation for writing the article (a response to what the author saw as an overreaction to Chinese models), but I feel the article goes too far the other direction, with multiple unsupported leaps of logic, and overstating the stickiness of AI client products.

  • The model that is most optimized around token cost is, in fact, Chinese. DeepSeek is astoundingly cheap by default, but if you use it from Reasonix (the harness optimized around its cache), it becomes even cheaper.

    • I keep seeing mention of the cache, what's special about it? All frontier llms have prefix caching, what is special about deepseek's approach?

      4 replies →

> Models are not free. Downloading them is free. Running them is not.

Is this really different from traditional software? Downloading postgres is free. Running it is not. You either buy hardware and assume the costs of owning and running that, or you pay to run it in the cloud.

  • I think the point here is that it takes the same hardware to inference an open source model as OpenAI/Anthropic inference their models.

    IE, a lower param OpenAI/Anthropic model can compete with a higher param open source model.

    So even if you are an American company who downloaded Chinese models in hopes of saving in cost, you still have to beat OpenAI and Anthropic in $/task which is very tough to do over the long run.

    • The problem is these AI companies are running at a loss with the hopes of pumping the prices after everyone is hooked. Now they are locked in at selling API access at rock bottom price. The valuations won't hold.

      1 reply →

  • But inference costs scale per task whereas platform services like postgres typically amortize across tasks. If you are selling tasks done by inference, then compute is part of your COGS and it scales per task.

    • Much more so, I would agree. A bit like how mainframe batch jobs might have once been charged back based on "cycles" used for a job or other tracked metrics.

The thing I do not understand here because it seems obvious: AI will be a commodity market and you simply cannot have a large PE multiple. So the valuations imagine a global commodity monopoly or duopoly coupled with the increased intelligence still disallowing other suppliers from becoming competitive? Without any network effects to help?

  • Perhaps to moneymen the difference between “ChatGPT” and the technology behind it isnt’t obvious. I’ve been very surprised at how few otherwise smart people are completely in the dark about how capable current models are.

    As soon as manufacturing starts building this stuff more, it will commoditize. The hardware prices won’t be terribly larger than the original. We’ll have a “Bambu labs” style company to make the AI OS, whatever that is.

  • There are many ways to create "sticky" products, even commodity products (e.g. coca cola, starbucks, etc). Network effects are just one.

    Regarding Price to Earnings... I'm not sure everyone fully fleshed out the end game for the frontier companies. It appears the pitch is that this current phase is a stepping stone to Artificial General Intelligence or Super Intelligence. I can understand the perspective of investors though... If you can get half of the world on these products at some point you will find something you can sell them even if its not the core product (i.e. loss leader).

  • AI (llm) will be a commodity market => I am not sure it was obvious. As of last month, folks thought open weight models are lagging by 6+ months. Once K3 is taken for a deep run across many use cases, it will be clear where it stands. But yes, I agree that now that intelligence is commodity, everything changes.

" - The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6."

This is a flatly false statement for most things powering backend applications. The AI consumer "doing real work" model, either for analysis, chat, or coding could well be more cost effective with closed frontier models.

But most of these internal glue business SaaS applications where engineers are integrating are not those tasks. It is those tasks which 1) drive immense amount of domain-specific data into the platform over time, and 2) are most encouraging of driving open model independence with no vendor lock-in.

Anyone on this site who has actually used ML models (more accurate in many cases) knows there's a lot of kludge that simply does not need a 5 minute agentic feedback loop to solve the problem. And they were solvable a year ago with lower class models. The token economics are exceptional and the anecdotes of a16z saying 80% of startups are productionizing open models is only surprising to people who think running your company on OracleDB in 2026 is a sound engineering decision.

  • You're correct, but that's a different market segment and not the market GLM 5.2 and its peers compete in.

    The labs are not interested in the small, fast, single purpose end of the market. Google increased their pricing on Flash so much that it stopped becoming a cheap model; instead, they released Gemma 4 open source, which is actually easier to use from a third-party inference provider than from Google.

    From a total token volume perspective, these "utility" models (classifiers, simple summarizers, small OCR models) will absolutely drive enormous volumes of tokens, at low prices and margin and modest overall market size. Because the models are small and the performance requirements are modest, and because their use cases are specialized rather than general, there are poor economies of scale: they can run cost effectively on rented small GPUs, and a big player doesn't get a structural cost advantage. These models are usually 1b - 30b in size, and can run on a rented 5090. I've productized these myself: I run millions of pages through a fine tuned 1b OCR language model that runs on 5090s at a cost far lower than commercial providers.

    But that's not the segment of the market where GLM 5.2, Kimi 3, etc., play. They compete with frontier capabilities, and they are not particularly cheaper than OpenAI models at a cost per task. (I do actually think they compete well with Anthropic, because Anthropic's model efficiencies are poor compared to OpenAI.) And although this part of the market may not be the bulk of the token volume, it is the bulk of the market value.

    That's because a lot of human knowledge work is too generalized and fuzzy for dedicated, fine-tuned models, so they are almost entirely different markets that don't particularly compete with each other. (Though if SaaS companies successfully build around verticals that can use small models applied against well-defined jobs, there may be opportunity to push the small/big capability boundary to subsume marginally more valuable tasks that today would require mid-grade reasoning.)

  • > "The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6." This is a flatly false statement.

    It may not be false but may be a "category error" [0]. Reserved GPU pricing & bulk inference pricing is 3x to 6x cheaper than "API rates", but renting your own GPU cluster (in this crunch) to run a 600b+ open weights is going to be "more expensive than GPT 5.6".

    Even then, it remains to be seen if Huawei will pull their weight (and match up to Nvidia) as spectacularly as their fellow Chinese AI Labs have. If so, the WAICO alliance is ready to go all-in.

    [0] Ben, and probably other "influencers" in this space, may be prone (knowingly or unknowingly) to favour points that make their conclusion for them (https://en.wikipedia.org/wiki/Motivated_reasoning).

    • Fair. Too strong a statement.

      But much like Ben's point that commoditization is a relatively novel concept to many in tech, it's not the consumer AI applications at risk of commoditization. They have distribution there.

      It's the literally millions of engineers who are updating codebases with tools replacing workers partially or wholly. It's the supply-side where there's compression, and no need for distribution.

      I would argue, given the enormity of the existing SaaS stack and how it integrates with the human machinery of personnel, that's where volume is. And that is clearly cheaper and a home run.

      Commoditizing a ~$100B AI consumer market is no small feat. Commoditizing 20% of the $500B SaaS market, to say nothing of the underlying systems in the who-knows-how-many trillions "Big Tech" market (you're obligated to say that like the Kool Aid man), is shocking.

> Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.

This is the story for Nvidia/AMD or cloud providers rather than OpenAI.

> With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates

It seems like there would be problems with this on both ends.

For general purpose models, everybody is trying to make them efficient, so you can't win just by being slightly more efficient. You would have to be so much more efficient that you can charge high margins while still capturing the majority of the market so that the high margins get multiplied by the majority of users and the users you leave on the table aren't funding open competitors. Meanwhile everyone else is also trying to improve efficiency, so one misstep and you're behind.

Example of where this can be a problem: You spend a preposterous amount of money to create an efficient model, then someone else publishes a paper with a new technique that gets a similar but incompatible efficiency improvement out of a model that costs a lot less to create. You have now spent an enormous amount of money in exchange for no competitive advantage.

And from the other end, one of the best ways to get efficiency is through specialization. A general purpose model can generate code or summarize a meeting transcript, but a special purpose model can do it as well or better with far fewer parameters and resources. But then you don't have a situation where one huge AI company has The Most Efficient Model, you instead have dozens of specialized models produced by independent sources that are each the best in a given niche. Any proportion of which could have open weights, or have an arbitrarily small advantage over the ones that are.

Moreover, these problems combine: Both the computing hardware vendors and the AI companies want the margin on doing inference, but the more of it one of them gets, the less the other does. If the AI companies were actually getting huge margins then it would be in the interests of Nvidia, AMD, Apple, Intel et al to fund efficient open weight models in the same way they fund Linux. Commoditize your complement. And those models don't even have to be better, as long as they're good enough that the closed models can't charge a significant premium and the margin shifts back to paying for hardware.

> for frontier work.

I'll agree that GPT 5.6 may well be the best given the above contstraint, but for run-of-the-mill dev tasks (real ones, not benchmark ones), GLM 5.2 still blows every other model out of the water.

Cost per task as a metric is a bit ridiculous because there are so many types of tasks. GPT-5.6 can do some tasks GLM could only dream of, but GLM can do some tasks 100x cheaper and better than GPT-5.6.

China is working on the whole supply chain though and they're willing to compete on razor thin margins. Just look at EVs. They build great cars but the competition is so aggressive that investing in any one Chinese EV company isn't exactly an amazing ROI.

I could see AI ending up the same way where the customer captures most of the value rather than the companies. Open weight models are what make that kind of competition possible.

> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use

I notice that the article, and this discussion, hasn't mentioned or considered local models.

We can already run a low-spec model on a laptop. Because there is demand for this, it will improve and we will get better laptops and better local models. We will also see models being run on dedicated local hardware and called from the laptop.

If I can download a reasonably capable model to my own hardware and run it without paying anyone for either the model or the inference tokens (effectively making models and intelligence actually free once the hardware is bought) how are the Frontier AI Labs going to make any money at all, let alone enough to support their vast valuations?

Yea but at those rates VCs will never make their money back. Because Deepseek and friends keep releasing the inference optimisations to everyone instead of holding them back to pay their investors.

  • [flagged]

    • This argument starts falling flat when AI companies literally stole from every business and country on Earth. China’s government can f right off, but so can this failed argument.

I did read the article, but it misses the core issue entirely, and it's why I shared my comment to begin with. Look at the cost-per-task benchmarks from Artificial Analysis https://artificialanalysis.ai/models?cost=cost-per-task

Anthropic’s API pricing is getting impossible to justify. Anthropic previously had the highest quality models, and used their position to charge premium prices, enjoying inference margins of over 70% [0]. They could charge these prices because no other model came close.

But over the past month, the market has shifted dramatically. Over every single performance tier, Anthropic is being squeezed on price.

* Low end: DeepSeek V4 Flash runs at ($0.02/task), Xiaomi's MiMo-V2.5-Pro at ($0.03), and Haiku at ($0.24). Anthropic is ~10x more expensive than the Chinese open-weight options.

* Mid tier: Claude Sonnet 5 ($1.53/task) is nearly 50% more expensive than GPT-5.6 Sol ($1.04), nearly 2x the cost of GPT-5.6 Terra ($0.82), and 3x the cost of GLM-5.2 Max ($0.47). There is basically no reason to ever use Sonnet 5, the competitors are significantly cheaper.

* High end: Opus 4.8 ($1.80/task) and Fable 5 ($2.75) are the two most expensive models, and GPT-5.6 Sol ($1.04) and Kimi K3 ($0.95) offer comparable performance for significantly less. Less the fact that Kimi K3 will get ~10x cheaper once its weights are released and served on neoclouds with Nvidia hardware [1].

OpenAI priced their latest GPT-5.6 models cheaply in order to regain market share. When Anthropic clearly had the best models, their 70%+ inference margins were defensible. But today they are the most expensive option in every single tier. Unless they make significant price cuts soon, they run a serious risk of bleeding market share.

[0] https://www.mindstudio.ai/blog/anthropic-inference-margins-7...

[1] "American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia hardware" https://x.com/rohanpaul_ai/status/2079027313455550839