Comment by jrflo
16 hours ago
The luna cost cuts were real though, not a one time promotion or something, due to some optimization (probably distillation?) that openai did.
16 hours ago
The luna cost cuts were real though, not a one time promotion or something, due to some optimization (probably distillation?) that openai did.
You assume that openai's inference is profitable and that they aren't just trying to bolster revenue before their IPO.
The only indication that openai is profitable comes from openai (whom I wouldn't trust with any statement, especially when it comes to profitability).
In fact there is evidence that inference is not profitable simply because the rate of losses doesn't seem to reduce as revenue increases: if inference had great margins, we would expect that as revenues increase, the amount of spend on training reduces as a fraction of total expenses. Since the loss-making fixed costs shrink as a fraction compared to the profitable inference, we should expect profitability to rise with total revenue.
However, all leaks of openai's numbers seem to suggest the opposite: as revenues increase so do the losses.
The indication that OpenAI's inference is profitable is that 3rd party providers host large models for cheaper.
Given that OpenAI is ahead in intelligence, it's also reasonably likely that they are at the frontier of efficiency too.
Your "evidence" for OpenAI's inference not being profitable is apparently based on leaked financials supposedly showing growing losses for reasons entirely unknown.
With their research, training, data centers, chip development, and hardware product development, there seem to be a number of reasons that might explain growing losses.
> Given that OpenAI is ahead in intelligence, it's also reasonably likely that they are at the frontier of efficiency too.
Frontier labs have no incentive to be at the frontier of efficiency.
Claude still leads the pack in general intelligence yet has the worst efficiency by far.
I don’t pay OpenAI’s bills - I pay what they charge me. Their cost accounting isn’t relevant to a user.
Argument was that open ai cannot be profitable with this. But sure, use it while you can.
6 replies →
what if it was because of quantization and they haven't released the new benchmarks for it?
Anything which changes the model needs new benchmarks I guess to compare with other models, otherwise you can benchmark Fable, and distill it to student model and keep claiming this is the Fable model
ARC Prize has retested Luna after the discount and validated identical performance.
(Also, quantization isn't inherently bad or damaging when done properly, e.g. QAT).
These APIs are used heavily by enterprises at scale; with lots of performance telemetry, live evals, etc. You can't really silently nerf API models at scale without people noticing.
Of course, what I said doesn't apply to non-API consumer sub models; there's many documented and officially confirmed instances of under-the-hood "juice/effort" adjustments. (Juice = a number your effort tier maps to underneath the hood; much like Inkling's effort=0.00 to 0.99).
Was it?
Given the timing, I think they A. shat their pants since Deepseek flash just came out with insane pricing before the price hikes, and B. Anthropic is really struggling in model tiers below opus.
It was smart for them to cut prices regardless of whether they had 80% efficiency gains or not