Comment by firasd
5 hours ago
One of the strangest things in the AI industry is 'tokenomics'. It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. This pattern has continued across various labs/providers for years--there is a continuous see-saw of pricing that doesn't seem related to anything.
So what open weight models do is at least provide a baseline of inference cost to add some sanity to the price markers. And of course predictability too--if you really want Kimi K2 instead of K3 you can still use it.
So the competitive pressure and predictability offered by open models is helpful for users
> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference.
The price is what the market is willing to bear for the available compute capacity and competitive landscape. You can only discover that price after trying different price points and seeing what happens.
Everyone is trying different pricing schemes and discounts as they test the market. The demand is fluctuating at the same time.
It’s probably very confusing if you’re primarily familiar with stable and mature markets. Price fluctuations are a common feature of new and evolving markets.
Most unmature markets aren't subsidized to the point that LLM market is, most market have some level of baseline profitablity, this market doesn't, that's because most market subsidized the marketing or the capex but this market doesn't hold the opex, the capex not the amount of marketing let alone all of this together
Most new markets are funded by initial investment capital. Early entrants operate at a loss as they grow.
This isn’t as unusual as some people are trying to make it sound. This has been happening since the dawn of finance.
I thought this would be less foreign to everyone since we just went through this whole conversation for a decade with Uber and Lyft. Their demise was predicted from the start from everyone who thought that it was going to collapse as soon as they couldn’t subsidize your rides with promos. There was much wailing and gnashing of teeth as their prices changed to feel out the market. Then they found profitability and the critics went silent.
3 replies →
[dead]
> if you really want Kimi K2 instead of K3 you can still use it.
I think this is a very important aspect, especially after the huge GPT-4o backlash when GTP-5 came out. Each model has certain quirks, and areas where the previous model might be better for some use cases than the latest and greatest, and the labs so far seem to have no desire to offer some kind of "LTS" release.
LTS is really useful framing even if increasingly distilled or slower on older hardware etc vs disappearing model acts
> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference.
FlashAttention was a hell of a drug.
Why do prescription medications cost so much, and generics so little (comparatively)?
Artificial inflation to recoup R&D.
Not sure pricing to recoup costs is artificial.
I think the argument is that artificial comes in with IP law, which some people feel is superfluous.
I do think that corporate price gouging is a huge problem that does need to be addressed. But especially with smaller business types — creatives, et al— I still haven’t gotten any grownup answers about what would compel people to get professionally good at something and innovate in the complete absence of copyright: the vastly better business model would be waiting for someone else to do something new and interesting, stealing their work, and then undercutting them in the market because you don’t have R&D/et al costs to recoup. You can’t say that wouldn’t happen because it’s exactly what the AI companies did to billions of people, scoffing at any protest. And ironically, they’re now whining about the Chinese doing it to them.
1 reply →
Separating out what is artificial or not seems more like an exercise in rhetoric; defining things as fundamental and real. The whole economy is an imaginary thing dreamed up by our natural human brains.
Not sure they're just recouping costs and not lining their execs pockets.
The government backed monopoly to ensure that supply remains artificially restricted to ensure that the market will support the higher prices is
9 replies →
What is strange about GPT-4 being expensive in 2023? Supply and demand. Which other model choices did we have? Prices are related to supply and demand. We see it play out with the introduction of capable open weight models or even other closed cloud models.
That’s a slightly naive view on pricing. That equilibrium point doesn’t just magically appear - it’s found through price testing.
Not really though right?
As of early June 2026, Opus 4.8 in fast mode cost $50/M output tokens and Opus 4.6 & 4.7 cost $150/M output tokens in fast mode
How can supply and demand explain the price drop? Was it cheaper to serve Opus 4.8? Is the demand for the newer Opus lower than for the older Opus? These are just fixed prices that seem picked out of thin air
They would have been picked out of thin air. That’s the joy of innovation - you have to randomly throw prices against the wall and see what sticks. The point where it sticks might be equilibrium or it may be an inefficient market… and nobody will know which one until it’s too late.
Did they actually ever cut the price on gpt 4? The oldest versions of it in the api still seem stupidly expensive? There were definitely price cuts as they introduced the turbo models and stuff, and new versions of each model might have gotten pricey cuts, but just because they're both called "gpt-4 something" doesn't mean they're the same under the hood or that they didn't change a bunch of stuff under the hood to make it cheaper to serve
> because they're both called "gpt-4 something
More like a generation of models with different specific use cases
Its very clear: nobody wanted to pay for usage at that price point
Quantization also started picking up around then, as well as distillation into smaller models
It's only strange if you do the silly thing of presuming a "fair market" in which e.g. it's generally easy to get reliable information about how all of the things work.
There's just obvious and enormous incentive for the OpenAI's of the world, along with all of the other players, to confuse, misrepresent or just straight up lie a whole bunch about everything given how new and unknown the tech is.
Yeah I've been thinking about this and the analogy I came up with is that tokens are basically equivalent to an in-game currency in -free to play" mobile games. You're trading actual money for some notional "curency" or coins that can only be used for one thing but, unlike mobile game coins, you don't know how many coins something costs before you use them. It's kinda weird.
I think it is worse actually. Tokens in a game are usually just purchased for enjoyment in the game. They are purchased as a part of your entertainment budget, not expected to be useful in any way.
LLMs are fundamentally tools intended to be useful. But LLM vendors don’t understand their systems well enough to actually price the product people are trying to buy (for example, the actual product of a coding model is the code that it produces, not the tokens, which are just an internal mechanical process involved in the creation of the code). Token based pricing is that lack of understanding leaking out of the organization that ought to be responsible for it, and being dropped on the user.
Imagine if we made cars like this! You’d go to the car dealer and ask for a car. They’d bring you a pile of parts, charge you for them, and try to put them together in front of you. You’d go back and forth for a bit, rephrase where you want the steering wheel, etc. Some of the parts wouldn’t fit but you’d be invited to pay for replacements as well. In the end you’d either have a car or not, that’s your problem.
Market discovery
Why is this so difficult to understand?
1. the field was nascent and new efficiencies were discovered
2. supply and demand
3. its in the company's incentive to make their models more efficient to increase overall usage so that while the margin remains the same, the total revenue + profit increases
I genuinely don't know what puzzles everyone?
Because as per usual it's silicon valley misunderstanding economics. AI is HPC. And how the HPC market worked before:
If you're the best performing "computing cluster" (ie. whatever you call the entity that can complete a massive calculation), you get a blank check from Congress.
Why? Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity. And of course, they were replicated worldwide for this reason. I mean not that anyone will admit this but we don't have the best possible solution, and we don't know either the upper or lower limits for fusion devices (plus the lower limit would be very useful for energy generation, which for the US would effectively mean almost literally unlimited large marine ships that never need refueling. And yes, the solution to that problem is almost literally a 3d shape. Not just that, but mostly)
Now it appears it does not work the same when you democratize computation. Humans want a particular amount of computation and are willing to pay a given price for that. But the accountants still saw the blank check from before and ... do what accountants do. Economics don't change because you make things bigger and accessible, do they? Oh ... wait a second ...
As someone put it recently though, we now have data. 2.3% of humans in the US are willing to pay $20 per month for the support of a model like GPT-5.5/Claude code. If that's true (and after years of having this model, why wouldn't it be?) ... it means AI startups are doomed (because it's not even 10% of what they need it to be to make economic sense).
> Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity.
We have a couple new nuclear weapon designs, but not really going for bigger or stronger. Just packaging.
We built the big powerful ones with 1960s computing.
Now, stockpile stewardship -- being sure that stuff will keep working without ongoing testing -- is a bit expensive in compute. You need early 2010s supercomputer power.
In other words, I strongly disagree that nuclear weapons are the primary driver of high-end compute.
> We have a couple new nuclear weapon designs, but not really going for bigger or stronger. Just packaging.
There's many other considerations. Like the type and amount of fissile material. To name one that became well known: any plutonium needs to be refreshed (re-breeded I believe is the term) every few years.
Also, look at first designs: https://www.bbc.com/news/newsbeat-35242069 The weight is secret, but I think you can easily see it's going to be deep into "extremely impractical" territory.
Surely you can see why someone (especially aircraft designers) might ask for better versions. Ideally you'd like a version that fits on the hypersonic missiles and those things ... are just not going to work. The size. The shape. The weight. None of them will work.
Then a quick theoretical exploration will tell you that the minimum theoretical size of such a device is tiny. The scare was about "suitcase sized", but if you actually do the calculation looking for the minimum ... to do it however you need to create an explosion of the correct shape to get anywhere near those minimum sizes. And explosion simulations are a problem that utterly sucks ... Oh and these are secret military projects, these simulations, not very optimal. The people doing them are best described as loyal, and not as great physicists. Not saying they're terrible, but in the movie Oppenheimer you can clearly see why the best and brightest are not available for these things.
High end compute also existed in the 60’s.
The military has already mastered fusion (power), and likely has mastered gravity in some form in secret. They aren’t using mainstream compute for these discoveries
1 reply →
I disagree with some of the framing, but that's what good discussion is about -- figuring out where we agree and disagree.
But consumer uptake strikes me as the worst way to judge whether the big AI shops will make it. That's not where most of the leveraged user return or deployable capital is.
I'm sure Teller could make you a 10 gigaton nuke with a slide rule if you did not mind some sub scale tests.
It's basically supply and demand?
That's not really a full explanation unless you have some idea about why supply or demand are going up and down so much.
Aggressive expansion of infrastructure, R&D and very volatile audiences.