Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the article says, so the threat is far from theoretical.
Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any hope at a successful IPO.
However the cold reality for both is that there is zero moat to a model anymore. It’s a pure commodity. Those selling compute and access to open models are gearing up to wipe the floor with Open AI and Anthropic.
The only moat lives at the Pareto frontier. If you are on the Pareto frontier you are good and can charge money. But the frontier is moving every week so it's super competitive. If you are the quickest innovator, I still believe there is a chance for a working business model for them
> However the cold reality for both is that there is zero moat to a model anymore.
The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability.
On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision.
On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places. In California or Europe this could be $30-60k per year.
As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less.
The last major moat is the ability for subscriptions to scale with need. This means easily adding and removing licenses. This is far easier than purchasing extremely expensive hardware (and managing it), and selling it if/when internal demand changes. It's the same reason companies use contractors. The ramp up/down costs are very high.
The only real moat that local LLMs have right now is privacy.
Hmm yes but the equipment cost moat is artificial. This scarcity was created by the big AIs by buying up all the future production capacity. That works for a while but it won't last forever.
It's the same with the subscriptions. Local models can't compete because they're simply giving too much value for money. They're effectively subsidised by Big AI. Again something that won't last.
You don't need to self host to get the benefits of an open model. There are many hosted providers cheaper than OpenAI or Anthropic who can give you a SLA, ZDR, BAA and all the other three letter acronyms your compliance department needs.
The important part is if they break the contract or raise their prices you can always move to a different provider. You get lower cost and lower risk at the same time which is extremely rare in business. That's just not possible for closed models where your only options are the official branded API or Azure/Bedrock.
OpenAI and Anthropic rent their compute from AWS & friends. When we say large enterprises are moving to open weight models it means they are cutting out the middleman and renting the compute directly from AWS instead of giving OpenAI and Anthropic a margin.
> they have at least a 3-6 month head start
This is a moat of nothing. Our company still hasn’t gotten access to Fable so switching to open weight models would mean getting access to similar quality models. In some orgs they are still on 2025 models.
"Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision."
Depreciation is designed to _encourage_ purchasing of useful local tools, by incrementally matching fractions of the cost of the tool to the revenue it generates over its useful life. The fact that a graphics card might have a book value of $0 after five years of depreciation is a feature, not a bug.
Since the invention of corporation tax it has also had the benefit of offsetting tax over the same period, instead of just one big offset in the first year.
Privacy is non-negotiable for corporate. Even without considering costs or country of origin, we've seen from OpenAI that claims of AI safety are worth less than the (virtual) paper they're printed on.
All it takes is one incident, and all your company's internal data will start showing up in public users' chats. You can rely on a contract to prevent this, or you can guarantee it by using a locally hosted model you fully control.
When combined with the cost savings and good enough performance mentioned in the article, this can become a huge selling point.
You’re missing the point that 98+% of the use cases for AI don’t require the latest greatest model and are far better positioned to use the fast-follow distilled cheap models.
OpenAI and Anthropic are fighting to win a race (build the biggest baddest model) that has no prize. The prize is mass adoption at scale at the best price, which is why companies are rapidly shifting to open model. They don’t need to pay 10x for a model that’s provides no practical additional benefit.
> As for intelligence, the frontier models from OpenAI and Anthropic are still superior
I'll grant they are superior at least right now. But also, they are too expensive.
We ($work) are finding that it is best to build engineering discipline around AI usage (who would've thought!) and use the cheaper models like Cursor Composer.
Using Opus we can blow through an entire month budget in an afternoon, so while more powerful, it is no longer practical except for rare very complex tasks.
> As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less.
I would argue that the reason they cost so little is because anyone can run open models and offer them as a service, so there's actual competition and the price is closer to cost. i.e. if the open models were just as intelligent as frontier models but cost the same to run as they do right now, the price wouldn't be higher (unless demand went up so high that marginal cost to provide more of the service went up, due to scarcity of hardware and or electricicy).
On the other hand, if what you're saying is the frontier labs have some pricing power due to their models being better, and that is the reason they are able to charge more than the companies providing open models as a service, then I would agree.
Actually, the moat is regulatory. Expect these companies to behave themselves in progressively more grotesque and sycophantic ways to get the federal government to make open/foreign models (and their output) illegal. After all, their very survival depends on it.
>On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription.
One of the basic questions/concerns here though is that it's not like the AI places are getting the GPUs for 10x less. It's true they have some economies of scale, but they also have some waste, and frankly in this particular case it's not clear they get that much gain over what a lot of businesses could achieve. The biggest traditional gain for central providers is that a lot of typical computing usage is burst-y, and in turn local kit might be underutilized. But with LLMs heavy users tend to use them all the time assuming their tokens allow it (and in the case of local hardware there's nothing stopping you, quite the contrary), they can use it directly interactively or leave them to go overnight on something too.
So it's reasonable to suspect that the reason subscriptions are only a fraction of the cost is that we're in a bubble seeing these companies losing money in an attempt to gain some sort of durable advantage. Just as every previous time, there is the chance that the music stops at some point, and they need to crank up pricing or pull other schemes to actually make money. Of course, it can be a good deal in the mean time, you basically get to suck down investor money for nothing, but it's also not unreasonable to at least be consider fallbacks. Even beyond questions of control and risk etc. I know at least a few places that are now genuinely considering questions like "what happens if a datacenter we depend on gets droned" that would have never had an iota of thought devoted to them even 5 years ago.
>On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places.
I don't think that's "surprising" at all, everyone knows about power use. And this seems like it gets heavily into what you're defining as "usable" and is also more useful to define in terms of cost-per-employee vs total. Obviously a bigger business will have a higher line number total even if the cost per employee is identical, but simultaneously can be expected to be making more revenue to pay for it.
If we're defining an average of a dedicated 5090 pulling 1 kW for every single employee (presumably some people wouldn't use it all the time, but others would then pull the compute for other work), running 24/7 (to cover people running stuff when they're away), then that'd be 8760 kWh per year. At my not particularly cheap New England location that'd be about $1900 per employee per year at the generalized residential rate (~$0.22/kWh), or $156 per month. That doesn't seem radical if it really does boost productivity. However, there is a lot of room to go lower. I'd expect a business to run backup anyway, and these days there are a lot of incentives to do that at least partially with batteries. That also opens up rate shifting as another way to pay back the cost. If we change to time of day pricing, that's 8 hours of peak pricing with the rest off-peak. 8 kWh of battery can now be had for a few thousand. And the off-peak rate is only ~$0.14/kWh, cutting the cost per year by about $700 to $1200 per employee per year. Solar power is also usually far more valuable to use yourself then sell back to the grid, and also continues to plummet in price.
None of this is to say that it makes sense for every place at all, but it's close enough to the the line that the math is at least worth exploring, or could at least lower the cost enough to be worth it given other things. It really comes down to how much extra value the company (or individual) expects to come out of it per month.
>In California or Europe this could be $30-60k per year.
Dunno about Europe, but at the kinda prices I see for California I'm really surprised more places aren't trying to move a lot of usage to battery+renewable.
>The only real moat that local LLMs have right now is privacy.
I don't think resiliency and control are things that can be taken for granted anymore, particularly on the global scale. War and terrorism is getting worse again. International relations are getting nastier, and governments have the power to just order places cut off. If LLMs aren't particularly valuable to a business, then why an expensive subscription? But if they are particularly valuable, then insurance is something leadership should be contemplating.
This is exactly true, I get annoyed by Claude one day and switch to something else, and the only thing that's ever keeping me tied towards Claude is the ability to search my old chats easily.
But Claude also makes it really hard to do that, so what am I even really paying for? Time to extract all my data, put it into a sqlite with FTS5 and make sure I never rely on the overly-opinionated, low-thinking PMs from these giant orgs again.
Of course, that "easy" step has lots of partial solutions like CTK (Conversation Toolkit) or MyChatArchive and I haven't found the perfect one yet, ideally it'd be something that dumped everything into Obsidian or an Obsidian-alike, but surely somebody is working on that? I'd pay $5/month for somebody to solve that problem for me, as long as I still owned the data...
Claude by default deletes old chats after a few months. I installed a custom end-of-session hook that throws chat transcripts into a database so that any model can read any other model's chat history. Super easy
Same experience, all my other friends in the industry report the same; at work we went from a huge push for ChatGPT last year to switching to Claude, back to ChatGPT when it became cheaper than Claude, and in parallel a deployment of open models being trialed with mechanisms to route to other models when needed (and based on pricing).
Over time I can imagine us becoming mostly open models on our deployments when hardware is more accessible and the need for expensive frontier models is constrained to very few use-cases that might demand their capabilities.
ChatGPT or ChatGPT Work or Codex? Are folks at your company still just using a chatbot to do work? (If so, this is quite surprising, as I find harness-based agent usage to be much much better than chatbotbot agent usage.)
I knew that there was no real moat from the very start, I mean, these things were close enough from the very start, how could it not result in a race to the bottom, especially as you can't really prevent distillation reliably?
I support and use open models as much as possible, but I'm not totally convinced that OAI or Anthropic have no moat, even as open models catch up to the frontier. Serving and inference are still hard problems when you're talking about a 2 trillion parameter model. Fine-tuning, if that remains a realistic need for businesses, is also a difficult infra problem at that scale. In the most bearish case, where there is no competitive advantage to using their models, big labs still have an advantage in this area.
Maybe there is some threshold where the price/quality math for your standard business tips in favor of smaller models and self-hosting the entire stack. I'd certainly love that.
A race to the bottom is where you lower standards, wages, or regulations to cut costs and attract business. What's actually happening is the opposite: a race to the top. Every model is trying to get better. Simultaneously they also happen to be getting more cost effective, but it's sort of a coincidence. Companies still require very good models, but they are not picking the "absolute best at any cost" anymore, because it turns out "any cost" isn't worth it.
Maybe open AI and Anthropic could just license their models to run on your own hardware. So a fixed cost instead of per token pricing or subscription with limits
This is the future, because the market will demand it.
I think there will be separate and huge markets for models, hardware and compute. That will maximize competition and innovation.
Why? Because even Blind Freddy can see the huge usefulness and power of these (and future non-LLM) models and no-one in their right mind is interested in becoming OpenAI's or Anthropic's bitch. Those companies have tickets on themselves.
Given the recent behavior of tech companies and the US administration, no one trusts either anymore.
The schadenfreude is that, at least for OpenAI, they were originally set up to make open models.
They were set up as a public benefit company and their returns were capped at 100x. They (well, Sam) went out of their way to put themselves in this death march to the IPO. If they had just done what Mark Zuckerberg did with Muse, they're not in this position.
Do you know how bad you have to be at the tech business to make Mark Zuckerberg look like a prudent-yet-visionary leader?
I know a few Australian devs who work at places also moving, or already have, from Anthropic… and not because of cost but because of Trump’s edicts to ban non-nationals using AI.
Think about how crazy it is for non-US companies to use American AI providers - their marketing boasts that you can treat their models as co-workers, assign tasks, invite them to slack annd video calls, etc. Taken at face value, would you hire someone remote who lived in a country that commonly does random shit like deciding whether or not remote workers aren’t allowed to go to work?
I know what happens in many big companies and not a single one is moving away from Anthropic/OpenAI/SpaceXAI.
> Unless they both dramatically slash prices then they’re in big trouble
False, they have already done so many times.
> Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any hope at a successful IPO.
False, margins are higher and I can have a formal bet that prices will go lower.
> However the cold reality for both is that there is zero moat to a model anymore
False, LLMs are not fungible and there exists a natural moat. I like the behaviour of Fable, not the behaviour of Opus - the fact that many people speak about this is evidence.
I don’t know what your sources are, but I work on AI at a Fortune 100 and open models now make up >90% of our internal token spend. Used to be 100% closed before this summer.
Not sure who you’re talking to but the NYTimes reporting clearly refutes your statements that nobody is doing it with clear facts. It’s been a tidal shift in attitudes over these last few months and the messaging back to OpenAI and Anthropic has been clear. Slash your prices by an order of magnitude or you’re done for most use cases.
We’re heading into corporate budget season for 2027 when all this is coming under a huge microscope in boardroom after boardroom across the country at a terrible time for companies trying to IPO.
> like the behaviour of Fable, not the behaviour of Opus - the fact that many people speak about this is evidence.
I'm wondering how much of that is the harness vs the model. Overall, the 'feel' of a model seems to be largely due to the harness than the model itself.
They kind of are fungible, up to a certain level of task. And much like most software developers don't need to exercise deep comp sci skills, most sw engineering doesn't have tasks that require the best models.
> I can have a formal bet that prices can go lower
Is this a typo? Have you actually made a formal bet on a prediction market or something to put your money where your mouth is, or are you just saying that you could? There's a lot of things I could plausibly make bets on, but that doesn't mean that they're likely to happen.
If companies are really doing this, then we're saying they have no problems spending tens of millions to get somewhat decent TPS and then having their employees complain they are timesliced and getting lots of timeouts because their org has 500 employees?
This is all spitballing, but I'd wager it's a third the type of workloads, a third hedging against your business depending on a single external provider, and a third trust.
Not everybody is coding or doing work that lends itself to burning tokens for warmth. Reuters for example seems to be more interested in using it for research, editing and formatting citations and the like. There's only so much of that work that needs doing, it doesn't always need to be real-time, and they probably don't see it scaling exponentially. They also need to be very aware and in control of their model's biases, or they risk it compromising their work output.
It's widely expected that all of the major providers will need to - and surely want to - drastically raise prices to justify the ludicrous amount of capital they're burning. Multiple companies have already talked about how their AI costs have exploded, and from what I understand that scale of enterprise is paying API rates. I would be disappointed if big business wasn't having a think about what that liability could look like. It's one thing to be reliant on a relatively "stable" vendor like Microsoft for Windows and Office, another to get AWS sticker shock, and then this is promising to be an order of magnitude worse.
Then just plain trust. What if ChatGPT starts recommending your competitors products, or the USA bars export of Anthropic's latest model (again, but for real this time), or they stop serving a model your business now depends on, and so on... That's a lot of risk to leave outside of your control.
I'm not sure why someone hasn't developed a company offering services that distributes AI across all idle or under-utilized VM's and PC's for enterprises in order to serve open sourced models. Outside of the electricity bill, there's no additional expenditure and you get the AI.
We've all seen the office spaces where there's 200 empty computers on a floor. Combined, it's something like 500 cores at ~3 Ghz each and around 3 TB of RAM. The networking is already there and software like exo already exists.
There are most definitely is a moat - but it works both ways. The railguards in the models create moats keeping customers out. And the cost to build a modern agentic model is in the 10 figure range and growing. This is an expensive arms race that is going to create moats.
But most commodities are the same way. It’s super expensive to drill for oil. I need oil and I’m in no position to mine my own because of the massive capital investment. But it doesn’t stop it from being a pure commodity.
I couldn’t care less which company drilled for the oil… it’s all the same to me. Models are increasingly no different.
OpenAI and Anthropic are a gas station saying “buy our gas for 10x the price!” When the world is looking at them saying it’s just gas, we’ll take the cheaper brand. We’ve tested your gas and it’s really no better than the stuff that’s 1/10th the price.
The open models are good because of distillation, which the US labs are actively working against via not revealing CoT ever and now you can see with OpenAI Astra 6 not even having a lot of CoT equivalents being emitted as tokens. Once the anti-distillation stuff is in place the open distillation models will probably start having larger and larger gaps.
If the companies survive the next few years, which they probably will because they represent too much of US economic growth to allow them to fail, this gap will keep on expanding.
Starting from zero without distillation is a lot harder, a lot more expensive and a lot more work. OSS models is what a laggard does to get adoption. China's gov't might keep on sponsoring it as a counter GPU embargo thing, but when gov't get involved, usually the other side gets involved too.
As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.
If we accept the premise that the top Chinese labs are simply distilling and can't compete otherwise: why don't US labs simply do the same thing? Distill their own models and slash their costs by 99% while keeping the same quality output. It should be a piece of cake if even the open labs can figure it out, after all.
One way or another they're getting the same results as proprietary labs, with a fraction of the hardware for a fraction of the cost. OpenAI can't keep raising funding rounds of $100billions to subsidize their compute costs and get results by brute forcing parameter count. And if they're having trouble keeping up with Chinese labs' efficiency, maybe they should stop worrying about distilling and instead hire some of the smart people responsible.
If that was the case the major labs would have done that.
They’re burning cash like there’s no tomorrow. They desperately need to show that they have a real business and not just a giant burning pile of cash doing academically interesting things. If they could simply sell models that are 95% as good at 1/10th the price they’d do that. They’re losing the enterprise sector because they’ve not done that.
Basically your argument relies on two claims, both of which must be true.
1. Competitors to OpenAI and Anthropic are good because of distillation.
2. OpenAI and Anthropic will come up with some methods for preventing distillation in the future.
Both of these are dubious imo. For RLVR tasks like coding in particular, you definitely don’t need continuous distillation to improve, otherwise OpenAI and Anthropic themselves would not be able to improve because there is no better model to distill from.
Yes, evidence is needed but especially for the claim that distillation is what makes these open models good. Serious citation needed.
Think about it: even if they distill the shit out of frontier models, the model still gotta learn, right?
If anything, as you can see from the K2 Horizon release, aggressive (self-proclaimed) reliance on distillation does not result in a model that has remotely any frontier capability. Try asking K2 Horizon to write iambic pentameter for instance, or even give it the car wash prompt. I tried both these on the Q8 quant for the 7B model and the results were depressing.
Open source does not apply to AI and we should discourage anyone from using that term. All the models are opaque and proprietary. You cannot go into any source code and fix bugs, or add features, or study it to learn more. It's not the same thing at all as open source software.
I swear, Qwen 3.8 27B @ Q8 is smarter than Sonnet 5 most of the time. Why wouldn’t corporate America self host at this point, especially with better options like Deepseek Flash and GLM 5.3 flash that’s a middle ground between Sonnet and Opus
Agreed. And conversely, American models can also just as easily be secretly influenced for bad things, or be more tightly controlled by the government, to corporate America's own detriment.
Still a ton of non tech companies doing a digital / tech transformation out there too lol. Maybe some so far behind they still have the real estate and rack space to get ahead on this one
The financials are incredible. If you spend one engineer's salary on hardware, you get a system running local AI that can multiply the efforts of an entire (small) team of engineers. It's a very "you can't afford not to" situation. Even considering the hardware prices today.
> Some U.S. firms remain reluctant to use Chinese A.I. models because of concerns over regulation and data privacy. AT&T researches Chinese models but is not using them, Mr. Markus said. Instead, it is working with popular alternatives made by American companies such as the Gemma A.I. model from Google and the Llama A.I. models from Meta.
This makes sense since corporations require legal certainty, and using an open model from an American company (probably) provides them some level of indemnity, and also someone to sue.
The current U.S. regime is also replacing some amount of that legal certainty with regime fealty. Picking Chinese options over American ones probably runs a risk of upsetting their leader. I've got to imagine American companies are weighing this factor in their decisions.
That's certainly one factor, yes. But even before getting to that part I think the bigger issues for Big Corp legal teams is mostly around the legal ambiguity of the models themselves. What representations are made about the training data? What jurisdiction governs the license? If somebody alleges that the model infringes their IP, what rights does AT&T have?
Counterparty risk is a lot more straight-forward to evaluate when dealing entirely within the US, with US companies.
Exactly. See TikTok trouble as example and quite honestly, try a local open source LLM and ask it to use profanity, paint nudes - the LLM doesn’t answer the question of it is from OpenAI or Google.
The thing is that needs more attention is reverse engineered a LLM which is highly fascinating. I tried it, but it seems I am not there yet to put it mildly. It requires serious effort.
I am just speculating but can LLMs be sleepers? You write software and it seeds traces here and there under certain conditions that pose a serious security risk.
Or a kill switch?
I don’t know. I distrust Chinese LLMs but even more due to training data.
It is after all not a Western model. Different biases and the might be subtle but nevertheless substantial.
In short: no open source LLM may be usable without additional Finetuning for certain valid use cases.
The real value is versioning and autonomy as well as lot more stable answering despite model rot.
Also testing and the supporting systems are easier to maintain.
Since most folks still somehow review the code, I find it highly improbable, it would become visible quickly. Unless say compiling Unix parts for example. But thats compiler work and not llm.
The license you receive when you download Gemma off of Google's website is not necessarily the same license that AT&T gets when they deploy Gemma as a customer service bot. That's the whole point. AT&T can work directly with Google for a licensing and legal framework that provides certainty.
Spent most of today reworking a rack and rig of gpus for all of our internal ai work… our big server is 8 rtx 6000 pro and 3 psu, I definitely feel I made a mistake not upgrading our wall power to 240v but so far we have multiple 30amp 120v and with 3 PSU uninterrupted power we have been very stable. My big upgrade will be moving to epyc motherboard from threadripper so we get gen5x8 with bifurcation instead of what we are stuck with today gen5x4 due to bios limitations . What has been so encouraging though is first deepseek v4 flash at 200+ t/s for single user and much more in aggregate- now on qwen 3.8 flash next for image support and eyes on glm5.3 flash for some testing … next is realistically considering co location and quiet a sizable loan to scale this to real hardware instead of miner rip vibes
What kind of cost are you looking at and how many people could use it? I’d be interested to know what the payback period is like, because Claude code is getting ridiculously expensive.
Bought the rtx 6000 pros one per month starting in January- since they doubled in price I wish I just used a line of credit back in Jan to buy all of them. For me it’s about keeping internal company content internal. Slack channels etc with tools for teams to use. Triage tools for Zendesk etc.
> AT&T turned to artificial intelligence models from Anthropic and OpenAI in recent years to help with customer service, call transcription and coding. [...]
> By May, open models accounted for 20 percent of AT&T’s A.I. use. That has since risen to 40 percent and may jump to 60 percent in the coming months, Mr. Markus said in an interview.
This is missing a crucial detail. We know they "help with customer service, call transcription and coding", but which of those have been upgrade to open models?
Call transcription is trivial to do with open models. I can run Whisper or Parakeet on a low-spec laptop.
"Customer service" could mean a lot of things, but it sounds feasible for open models too.
"Coding" - they might go to open models for that, but I expect the costs involved in paying for closed models for software developers within AT&T are a fraction of the costs involved in transcribing all of their calls or handling aspects of custom service for millions of customers.
From later in the story:
> AT&T researches Chinese models but is not using them, Mr. Markus said. Instead, it is working with popular alternatives made by American companies such as the Gemma A.I. model from Google and the Llama A.I. models from Meta.
I'd assume the author is just getting confused because of ollama and llama.cpp and all the other ecosystem "llama" that are still in use. Llama really did kick off the open models thing
I'd love to! For real coding though, SOTA models barely get the job done. It wasn't until Opus 4.5 that you could really get decent results.
I'm sure this will change (and I can't wait for it!) but as of today, open models might be fine for summarizing and writing docs, but you need SOTA to work on code if you want to be competitive.
Every time I check in on this I hear a more recent model is the one where they started doing good work. I'm excited to here that Astra is where it got capable enough to work on code next year
I do think there are now open weight models that are on par with (or beating) Opus 4.5 by now (e.g. Kimi K3, GLM5.3). But yeah obviously the frontier closed source models seem to have pulled away once again, so open weight seems to be a few months behind right now (which might be too long to wait for a lot of people!).
You're not corporate America (and trust me, I mostly mean that as a plus).
I also work in software, and while I vaguely disagree that open models can't be used (they absolutely fit into productive niches here, and holy hell are the last generation [ex laguna s1, kimi k3, glm 5.3, etc] actually decent) - I will agree that SOTA are a better fit for software development, especially when used in conjunction with an already very expensive employee who's driving them.
But for "Corporate America"... no. You absolutely don't need SOTA. They're doing things like transcription, summarization, customer interaction, minor technical tasks like form creation in existing tools, report generation (ex - powerpoint, pdf, docs, etc) and other general "white collar tasks". Think about roles in business that are in the 60-85k compensation range.
It's mostly busy work that keeps existing processes flowing and the business on the rails. Important, but not research/novel.
And cheap ai... is a wonderful fit for a lot of this. No one wants to replace an employee making 80k with a less reliable AI that costs 45k a year in tokens (SOTA). But they're absolutely willing to drop 2-3k/year on AI (~100/month - right in the open model cost range) for that employee if they can get a 10% bump in productivity or happiness.
Greg Kroah-Hartman from the Linux kernel team said recently in a talk (https://www.youtube.com/watch?v=_MwMLPmMccs) running open models locally is already good enough for finding Linux kernel bugs and generating patches to fix them. He said that the false positive rate he's encountering was only one third to one quarter IIRC. He doesn't provide any details but it didn't sound like he was running some kind of huge server or something either, just a desktop workstation of some sort.
"real coding" carries a lot of the weight in that comment.
Seems like if you ask 5 different people what "real coding" means you might get 5 different answers.
Not everyone is building the next framework or compiler.
Self-hosted Qwen 3.8 @Q4 on my RTX 3090 can produce beautiful functional CRUD pages and apps all day long. And that is 90% of the "real coding" being done in corporate settings.
The quote in the article about Mazda vs. Maserati captures this. Many might want the Maserati and drool over its specs and capabilities, but balk at the cost and how often are they really going to run it up to full performance limits on their daily commute to their cubicle?
Yeah, I totally agree. I'm sure these open models are more than good enough for non-dev work. I'm also sure they'll be good enough for dev work soon enough (and some people are saying the latest already are). My point was that given the difficulty of writing great code and dealing with large systems, SOTA just recently emerged as a viable option. I expect open models to catch up soon.
I'm not seeing it. Corporate America needs someone they can sue if anything goes sideways with AI given the rate of change and legal ambiguities. It took years/decades for actual, real, open source to be widely adopted in corporations for the same reasons.
The main indemnity that providers like GitHub Copilot provide is against being sued by the copyright holders whoes works were "transformed" to make the models. They generally don't offer that for the Chinese models, at least in my region. I can't see any reason other than back-room deals making it so: they're all using stolen works, and all could spit out memorized fragments. But if you're talking specifically about the fitness or otherwise of the actual software these models produce, that is generally handled with the use of sacrificial employees.
I had a conversation with a large group of friends and we independently came to the conclusion that Openai/Claude does not deliver more than a open source model. It takes about the same and the quality is about the same, and this does not mean it is good
I'm just glad the dweebs who were parroting "OMG, Distillation attacks!!1!" have been empirically proven wrong, and discussion around open-weight models are more rational now.
The thing is, for big companies (or even small ones owned by PE, which is MOST of them), it's not just about cost. The big thing is risk.
In my experience as a tech diligence assessor for PE firms for the last 7 years, investors really, really don't like companies being beholded to single entities that they don't control. Anthropic and OpenAI have demonstrated that they are not trustworthy, or predicatable, or finanically safe, or even capable of hitting three fucking nines. Investors know they need companies to be on the AI train, but they really don't like vendor lockin to the big AI companies. Every diligence I get asked "how easily can they change models?"
I think when open models reach 80% or 90% capability (or maybe even less!) a whole lot of companies are going to say "almost as good with way less risk is a better deal".
Well the big thing that stands out in the reliability front is that AWS and Google cloud are typically stable, certainly more than 3 nines. Meanwhile openai seems to roll dice on every request to see if they’re going to return a 500 or not.
IMO this will be a blip. There’s a lot of talk in the wake of all the Uber handwringing about token spend. Legacy enterprises want to look innovative to Wall Street without spooking them, so it’s easy to hop on the narrative and “show” that they’re innovating in a cost responsible manner.
This feels reminiscent of the big push to RAG a few years ago. And, more broadly the skunkworks projects that big companies tout in the press before they end up killing, when the operational overhead becomes too much for their liking.
Ultimately, the narrative is good for the consumer and the enterprise. It’ll mean OpenAI and anthropic will have to keep prices low. But ultimately, in the course of the next 10 years, I don’t see enterprises wanting to do this themselves. It’ll just be simpler (and eventually safer in their eyes) to send traffic to the big labs.
I don't see this being a blip, for multiple reasons.
A) the uptime matrix between Github, OpenAI, and Anthropic means that we've faced multiple entire days of not being able to ship org-wide due to our reliance on automated code review and other tooling. Every cloud service baked into our CI pipeline becomes a point of failure. We can't live without AI anymore, but it too often either directly or indirectly gets interferes with our ability to ship, and I don't see this improving any time soon.
B) Locally hosted AI has serious advantages with regard to PII/sensitive data management, and there's not much the frontier models can do to overcome this. There are so many things I want to build and let loose in a sensitive data environment but can't due to data governance around frontier models.
C) Anthropic and OpenAI cannot keep prices low forever. They're still burning insane amounts of cash and at some point, they're going to have to transition from growth mode to profit mode. They're already juicing their sales pipelines to the max with introductory pricing and other things to get people in the door. But those are all short-term online marketing plays.
I think ultimately the folks with the purses won’t care enough about a for it to be taken seriously, even if it’s an engineering bottleneck.
B definitely has scope but still smaller than I’d expect. When I was an intern at yelp, I was migrating us off internal credit card management to Braintree/stripe. No one would have imagined outsourcing that in early 2000s. There’s a long tail of stuff that you can’t sent to a 3rd party; but for most use cases it’ll suffice.
For c, true; but this is expensive for everyone, including the Chinese model companies that are trying to undercut Open Ai / anthropic. In the nth degree, i think the field will bring the cost down to the place where it’s manageable (see the existence of the cheap Chinese models). The real question is the $$ spent on pushing the research forward at scale.
A lot of corporate use doesn’t require SOTA models. Things like summaries or writing comments in Slack are light work. For those reasons, local AI makes a lot of sense for some companies.
The article clearly distinguishes between open source and open weights models, and the headline and parts of the article state that it's open source models that are having a moment. But it doesn't list any examples. The only model families explicitly named are open weights only.
Could someone clarify whether there are actual open source models that are competitive with the likes of Gemma (mentioned in the article), or is the headline just wrong?
I use opensource models at work because my work is too cheap to spring for a $20/mo account for me. Since HuggingFace models can be run on my laptop now (still very slow though), nothing is leaving the 'secure environment' and so I can actually get work done (instead of the 'old' version of coding and writing - google).
You still need the models to be able to perform web searches, don't you? In which case the data goes in and out of your machine and there is risk for prompt injection attacks. I think it's needed at least for documentation purposes.
Been wishing my company was doing this. We are locked in with chatgpt for chat and claude for coding. I have a 36gb m4 max laptop but I'm not allowed to run anything on it even though it would be totally free. Absolute pain.
Easy prediction: LLMs will get shrunk down further and further until GenAI is just something that ships on a chip as part of your hardware. In the future it will seem quaint that we needed a network connection to talk to our LLM.
Adoption of open-source models to my mind is a similar step in that direction. In all cases, the goal is to become untethered from a mercurial vendor.
Like taalas.com (very recently acquired by AMD), or cerebras.ai (whole wafer is a chip)? As you said, I also think that is one of the main direction many companies (and academia) is moving to.
We're gonna start baking in models like TTS with thousands of voices available in any language as a chip on device. They just need to hit 99% accuracy and then it's a done deal.
I currently have a small TTS model running in the background on my machine through which my agent(s) speak to me as they work. If that can be baked into an ASIC along with a few thousand voices in every major language then it should just be a utility chip on your mobo for anything that needs it. And yes, I too, am looking forward to it.
The amount of downtime the big AI companies are showing doesn't help. Do you want your company's sales and customer service to shut down every time OpenAI or Anthropic goes down?
This is a prescient concern and it all came about due to the initial export controls placed on Fable on a Friday afternoon with zero warning. It was partly Anthropic’s fault though. Engaging in criti-hype by going on an advertising blitz explaining how their new model is some kind of genius super hacker didn’t help their case.
Nevertheless, it happened. The problem is that businesses demand stability. Building an application, service, or business process around an API that can be turned off whenever a government demands represents an unacceptable risk.
It’s not a mental exercise when it has already happened once before and likely will happen again. If you control the weights and invest in your own hardware to run them (or rent it from a cloud provider), that risk can be mitigated.
I actually think they’re hooked on models they can fine tune.
You can’t further train the closed models. The open models can be fine tuned for your company. Big companies fine tune models on all the internal systems and documentation, not just through .md files (you’d blow up the context trying it that way) but actual fine tuning of open weights models. A low tier but open weights model actually beats frontier models when you do this for a specific task.
I think the frontier providers need to have a way to isolate instances (bedrock style?) and allow fine tuning to compete. Big companies are absolutely fine tuning models right now and getting better results than even the best frontier models for their use cases.
What IDE/extensions do you use for open-source LLMs? I tried VSCode with ollama and lm studio, and the experience is very subpar to the built-in copilot. It's not very usable.
Seconding the TUI usage. Opencode, codex, and pi are ones that I have / my friends have had success with using locally. I mostly use pi. It feels like the most "boring tool that just does its job" out of the big options.
I’ve been using Opencode, but recently migrated to oh-my-pi and have really enjoyed it.
I’ve also had luck with copilot-cli, but find that to be more limiting and it’s only really worth using if you are paying for GitHub Copilot (or your company pays for it, as is my case).
At our small company we are hooked on individual subs. But yeah a larger dev shop can't really pull that off and I get how they'd be dying by the token cost.
I have recently come to the conclusion that thinking for 2 seconds and using a cheap model with a slightly more detailed prompt works just as well as zero-shotting an idea with a fancy model. I work in science, and instead of asking the model “write a topic extraction algorithm”, I just say “hey look at this matrix factorization script I found in a repo, now make it use plotly and duckdb”. Have others come to the same conclusion here?
It makes me skeptical that the flagship companies are sustainable. Every company is going to maximize “fuel efficiency” to save time and money.
Then again, maybe the cheaper models have more markup for them, in which case they are probably happy w this arrangement. I’d be curious to know how the money making varies by model.
Really? It's worth it to you to spend 10 minutes thinking about how to prompt a dumber model to save $0.05? (not that open source models are dumber any more)
Its not just about saving $0.05. There are many legitimate reasons to not want the AI to do the thinking (aka architecting or planning) for you. In those cases where one-shotting is undesirable or unnecessary, closed frontier AIaaS has no advantages over open weights.
Well its def not $0.05, I just started using claude sonnet 5 and Ive found most simple questions might be 0.05 cents, a unit test is something like 0.10 -> 0.20 and small features and classes get into the individual dollars. Sure its a lot faster but at the end of the day its not cheap.
Plus there is something to say about being in the drivers seat, youll have a much better idea of how it works instead of needing to talk to claude and hope its correct. Since most LLMs also not very good at ideas even in my experience with better models its better to think for 10mins, youll get a much high quality result
I think using open-source AI is no longer about API cost but about company survival.
Take Anthropic for an example. Anthropic has successfully destroyed customer trust, at least for me. DHH in a recent interview mentioned that Claude refused to translate an article about immigration. Not summarize. Not editorialize. Translate! I think this reveals an unacceptable level of paternalism: Anthropic fundamentally believes that it possesses a moral authority superior to the people actually paying for the API. If such basic and mechanical translation is already too sensitive to touch, the goalposts have moved from safety into outright censorship. What prevents them from quietly deciding tomorrow that your proprietary business logic, financial data, or legal documents cross their invisible moral line?
Let alone how Anthropic treats Cursor and Figma - not that they are wrong as companies are free to compete legally, but nonetheless it shows that companies can't outsource their intelligence to a potential competitor.
I get what you're saying and it's concerning how much power these big labs have amassed and how little transparency there is in what they do with it...
But I doubt this a major factor in the trend. I just don't think it's something most corporate users run into. My understanding is these guardrails are negotiable for enterprise customers anyway.
And, not for nothing, but if I owned a human-powered translation company I would've refused to translate it too.
I like the Claude constitution overall - I hope it becomes something representatives vote on and amend, to avoid the centralized corporate censorship you describe. In the meantime, I am fine with it abstaining from doing DHH’s bidding, especially because there are so many AI alternatives.
Ah, yes, I'm sure the article that moral paragon DHH wished to translate was not at all harmful, and that this was a good-faith effort on his part /s
While I agree that Claude can be overly paternalistic at times, how should it respond to a request to translate, say, bomb-making instructions? It's reasonable to me that it might refuse this.
The innovative edge markup already faded and the race is to the bottom, more features, more reach, less cost. It's going to be extremely hard to recoup those giant investments. No, the bubble won't pop, it already popped and morphed at the speed of AI that we didn't even notice, money just realigned, llms keep pushing the frontier, and peripherals are gaining momentum
Most open source fans are also hostile to copyrights existence and are openly IP abolitionists. As such, they collectively respond with "good."
This is actual communism, and the fact that Bernie Sanders and every other member of the DSA isn't actively fighting for open source and is often fighting against all AI shows how fake their purported movements are and have always been.
If IP were entirely abolished tomorrow with no other change to our economic system, you would still not have universal Healthcare (maybe drug prices would be lower, though), Elon Musk could still donate however much he wants to get his preferred politicians elected, fossil fuels would still be used in amounts that destroy the world, etc. Very importantly, AI would still be used to try to manipulate and control the public.
The fact they prioritize other fights more than OSS, and have a rather dim view of AI, is hardly proof that they are fake.
> Most open source fans are also hostile to copyrights existence and are openly IP abolitionists.
Open source licenses are only enforceable because of copyright law. How are you going to enforce GPL3 when you have no legal authority to say what people are allowed to do with your code?
Who cares about power user fans? They don't own the copyright.
Open source authors have always been protective of their copyright. There are numerous examples when drivers have been copied between BSD/Linux (I forget which direction) which led to huge flame wars.
The whole point of the GPL is that it uses copyright and copyright assignment to the FSF to protect what it calls software freedom.
BSD authors are very upset if the attribution clause isn't observed. And so on.
It is communism to exploit poor open source authors? I have to read Marx again.
Anti-capitalism does not automatically qualify as communism. Notably, if IP were to be abolished, it doesn't belong to anyone. The expectation that open anything includes some sort of DRM-like content gating is counterproductive.
It's possible that LLMs will be a commodity in the future. Just like airline industry, AI will be tremendously important for society, but AI companies will not be making lots of money. It will be Nvidia, Micron, Dell and others shovel makers making money.
My hot take is that open models don't really save you money and introduce more router complexity and security risk (because you're now sending your company data through more less trustworthy providers). Look at cost per task not cost per token and the pareto curve is largely owned by closed models.
Just use Fable 5.1/Opus max for the hardest problems, GPT Sol high as your workhorse, and maybe terra for async batch stuff you don't really care about. Gemini 3.8 High also looks pretty good and is quite fast if you're already a GCP shop. You can basically benefit from open models without using them because they force the frontier models to be cheaper.
We run Gemini fast for random end user queries for general staff.
We have our own on-premise inference server (quad MI300A) that runs Kimi 2.8 extremely well and we transitioned all heavy work to it since it's basically instantaneous for the whole team. It's a good enough solution and we will hit break even before the end of the year already.
Not everyone needs frontier models and availability is frequently much more important than a lot of companies realize.
This is not a viable strategy. You're effectively paying for hardware and/or cloud. Why do this when OpenAI and Anthropic are both significantly subsidizing costs to win the market?
Talk to “AI” executives at large firms and 95% of them are clueless sales types that crawled their way to the top. Then again, it is basically a repeat of IBM, Microsoft, Oracle, etc. Same dumb executives making decision to not get fired and enjoy their place at corp.
Some of the most insidious parts of AI infrastructure includes the embedding model. Corporations have already spent an outstanding amount of time and money creating embedding vectors that are closed source and not reproducible. This means that all their data is locked into whatever embedding model they chose initially.
I highly recommend utilizing an open sourced embedding model instead of paying for a closed source one. It's vastly more reasonable to run an open sourced embedding model as a first step. They're much, much smaller and, due to the overhead of network latency, and running it locally has almost the same speed as through an API even on slow computers.
I would even go so far as to say that closed source embedding models have a high risk of data hostage. If a team doesn't have access to the embedding model, the embeddings become useless. A corporation like OpenAI could, say, hike the prices to that model by 1000x and everyone would have to pay up or forfeit any utility of the data.
I envision a future where open source embedding models are shipped with relevant technologies and implemented by currently under-utilized chips like NPU's. A startup developing cheap microprocessors that can run them is an idea I would pay cash for. Or perhaps they will be bundled with security tokens.
While it might be impractical for all corporate teams to run language models, it is very realistic for everyone to operate an open sourced embedding model, at least in their private cloud. Better yet, utilize transfer learning on an open sourced one to train your own, that way the embedding vector is more secure against competitors and trade secrets.
What are people even using embeddings for these days? It certainly seems like giving an agent grep covers most of the use cases. Dare I say: grep is all you need.
how much of performance comes from inference time tricks like scaling, topn ect .
maybe models providers are also in position to run their models vs running os models by a generic providerc
Meanwhile the rest of the world tries to un-hook itself from corporate America. Too many problems coming from the USA lately - it is not worth it to help sustain this anylonger. Canadians have realised this - others are realising this as well right now. Mr. Trump "no more forever wars", starting another forever war.
Using cloud hosted models that are created as a psy-op by domestic billionaires and crypto-fascists will end badly.
At least with self-hosted models the people running them get full control and don’t need to worry about the model or guardrails changing under their feet.
Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the article says, so the threat is far from theoretical.
Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any hope at a successful IPO.
However the cold reality for both is that there is zero moat to a model anymore. It’s a pure commodity. Those selling compute and access to open models are gearing up to wipe the floor with Open AI and Anthropic.
The only moat lives at the Pareto frontier. If you are on the Pareto frontier you are good and can charge money. But the frontier is moving every week so it's super competitive. If you are the quickest innovator, I still believe there is a chance for a working business model for them
> However the cold reality for both is that there is zero moat to a model anymore.
The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability.
On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision.
On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places. In California or Europe this could be $30-60k per year.
As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less.
The last major moat is the ability for subscriptions to scale with need. This means easily adding and removing licenses. This is far easier than purchasing extremely expensive hardware (and managing it), and selling it if/when internal demand changes. It's the same reason companies use contractors. The ramp up/down costs are very high.
The only real moat that local LLMs have right now is privacy.
Hmm yes but the equipment cost moat is artificial. This scarcity was created by the big AIs by buying up all the future production capacity. That works for a while but it won't last forever.
It's the same with the subscriptions. Local models can't compete because they're simply giving too much value for money. They're effectively subsidised by Big AI. Again something that won't last.
You don't need to self host to get the benefits of an open model. There are many hosted providers cheaper than OpenAI or Anthropic who can give you a SLA, ZDR, BAA and all the other three letter acronyms your compliance department needs.
The important part is if they break the contract or raise their prices you can always move to a different provider. You get lower cost and lower risk at the same time which is extremely rare in business. That's just not possible for closed models where your only options are the official branded API or Azure/Bedrock.
I think you went from one extreme to another.
OpenAI and Anthropic rent their compute from AWS & friends. When we say large enterprises are moving to open weight models it means they are cutting out the middleman and renting the compute directly from AWS instead of giving OpenAI and Anthropic a margin.
> they have at least a 3-6 month head start
This is a moat of nothing. Our company still hasn’t gotten access to Fable so switching to open weight models would mean getting access to similar quality models. In some orgs they are still on 2025 models.
"Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision."
Depreciation is designed to _encourage_ purchasing of useful local tools, by incrementally matching fractions of the cost of the tool to the revenue it generates over its useful life. The fact that a graphics card might have a book value of $0 after five years of depreciation is a feature, not a bug.
Since the invention of corporation tax it has also had the benefit of offsetting tax over the same period, instead of just one big offset in the first year.
1 reply →
Privacy is non-negotiable for corporate. Even without considering costs or country of origin, we've seen from OpenAI that claims of AI safety are worth less than the (virtual) paper they're printed on.
All it takes is one incident, and all your company's internal data will start showing up in public users' chats. You can rely on a contract to prevent this, or you can guarantee it by using a locally hosted model you fully control.
When combined with the cost savings and good enough performance mentioned in the article, this can become a huge selling point.
8 replies →
You’re missing the point that 98+% of the use cases for AI don’t require the latest greatest model and are far better positioned to use the fast-follow distilled cheap models.
OpenAI and Anthropic are fighting to win a race (build the biggest baddest model) that has no prize. The prize is mass adoption at scale at the best price, which is why companies are rapidly shifting to open model. They don’t need to pay 10x for a model that’s provides no practical additional benefit.
3 replies →
> As for intelligence, the frontier models from OpenAI and Anthropic are still superior
I'll grant they are superior at least right now. But also, they are too expensive.
We ($work) are finding that it is best to build engineering discipline around AI usage (who would've thought!) and use the cheaper models like Cursor Composer.
Using Opus we can blow through an entire month budget in an afternoon, so while more powerful, it is no longer practical except for rare very complex tasks.
Re this point
> As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less.
I would argue that the reason they cost so little is because anyone can run open models and offer them as a service, so there's actual competition and the price is closer to cost. i.e. if the open models were just as intelligent as frontier models but cost the same to run as they do right now, the price wouldn't be higher (unless demand went up so high that marginal cost to provide more of the service went up, due to scarcity of hardware and or electricicy).
On the other hand, if what you're saying is the frontier labs have some pricing power due to their models being better, and that is the reason they are able to charge more than the companies providing open models as a service, then I would agree.
Actually, the moat is regulatory. Expect these companies to behave themselves in progressively more grotesque and sycophantic ways to get the federal government to make open/foreign models (and their output) illegal. After all, their very survival depends on it.
>On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription.
One of the basic questions/concerns here though is that it's not like the AI places are getting the GPUs for 10x less. It's true they have some economies of scale, but they also have some waste, and frankly in this particular case it's not clear they get that much gain over what a lot of businesses could achieve. The biggest traditional gain for central providers is that a lot of typical computing usage is burst-y, and in turn local kit might be underutilized. But with LLMs heavy users tend to use them all the time assuming their tokens allow it (and in the case of local hardware there's nothing stopping you, quite the contrary), they can use it directly interactively or leave them to go overnight on something too.
So it's reasonable to suspect that the reason subscriptions are only a fraction of the cost is that we're in a bubble seeing these companies losing money in an attempt to gain some sort of durable advantage. Just as every previous time, there is the chance that the music stops at some point, and they need to crank up pricing or pull other schemes to actually make money. Of course, it can be a good deal in the mean time, you basically get to suck down investor money for nothing, but it's also not unreasonable to at least be consider fallbacks. Even beyond questions of control and risk etc. I know at least a few places that are now genuinely considering questions like "what happens if a datacenter we depend on gets droned" that would have never had an iota of thought devoted to them even 5 years ago.
>On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places.
I don't think that's "surprising" at all, everyone knows about power use. And this seems like it gets heavily into what you're defining as "usable" and is also more useful to define in terms of cost-per-employee vs total. Obviously a bigger business will have a higher line number total even if the cost per employee is identical, but simultaneously can be expected to be making more revenue to pay for it.
If we're defining an average of a dedicated 5090 pulling 1 kW for every single employee (presumably some people wouldn't use it all the time, but others would then pull the compute for other work), running 24/7 (to cover people running stuff when they're away), then that'd be 8760 kWh per year. At my not particularly cheap New England location that'd be about $1900 per employee per year at the generalized residential rate (~$0.22/kWh), or $156 per month. That doesn't seem radical if it really does boost productivity. However, there is a lot of room to go lower. I'd expect a business to run backup anyway, and these days there are a lot of incentives to do that at least partially with batteries. That also opens up rate shifting as another way to pay back the cost. If we change to time of day pricing, that's 8 hours of peak pricing with the rest off-peak. 8 kWh of battery can now be had for a few thousand. And the off-peak rate is only ~$0.14/kWh, cutting the cost per year by about $700 to $1200 per employee per year. Solar power is also usually far more valuable to use yourself then sell back to the grid, and also continues to plummet in price.
None of this is to say that it makes sense for every place at all, but it's close enough to the the line that the math is at least worth exploring, or could at least lower the cost enough to be worth it given other things. It really comes down to how much extra value the company (or individual) expects to come out of it per month.
>In California or Europe this could be $30-60k per year.
Dunno about Europe, but at the kinda prices I see for California I'm really surprised more places aren't trying to move a lot of usage to battery+renewable.
>The only real moat that local LLMs have right now is privacy.
I don't think resiliency and control are things that can be taken for granted anymore, particularly on the global scale. War and terrorism is getting worse again. International relations are getting nastier, and governments have the power to just order places cut off. If LLMs aren't particularly valuable to a business, then why an expensive subscription? But if they are particularly valuable, then insurance is something leadership should be contemplating.
This is exactly true, I get annoyed by Claude one day and switch to something else, and the only thing that's ever keeping me tied towards Claude is the ability to search my old chats easily.
But Claude also makes it really hard to do that, so what am I even really paying for? Time to extract all my data, put it into a sqlite with FTS5 and make sure I never rely on the overly-opinionated, low-thinking PMs from these giant orgs again.
Of course, that "easy" step has lots of partial solutions like CTK (Conversation Toolkit) or MyChatArchive and I haven't found the perfect one yet, ideally it'd be something that dumped everything into Obsidian or an Obsidian-alike, but surely somebody is working on that? I'd pay $5/month for somebody to solve that problem for me, as long as I still owned the data...
Claude by default deletes old chats after a few months. I installed a custom end-of-session hook that throws chat transcripts into a database so that any model can read any other model's chat history. Super easy
> search my old chats easily
Don't rely on chat history. Have it write and maintain summary files that you can import into different sessions, at least for anything important.
Just use a proxy and log everything.
1 reply →
> the ability to search my old chats easily.
I’d try:
1 exporting my data (I imagine it’s common outside of GDPR?)
2 asking Claude to convert it to an easily digestible format :)
Same experience, all my other friends in the industry report the same; at work we went from a huge push for ChatGPT last year to switching to Claude, back to ChatGPT when it became cheaper than Claude, and in parallel a deployment of open models being trialed with mechanisms to route to other models when needed (and based on pricing).
Over time I can imagine us becoming mostly open models on our deployments when hardware is more accessible and the need for expensive frontier models is constrained to very few use-cases that might demand their capabilities.
ChatGPT or ChatGPT Work or Codex? Are folks at your company still just using a chatbot to do work? (If so, this is quite surprising, as I find harness-based agent usage to be much much better than chatbotbot agent usage.)
2 replies →
I knew that there was no real moat from the very start, I mean, these things were close enough from the very start, how could it not result in a race to the bottom, especially as you can't really prevent distillation reliably?
I support and use open models as much as possible, but I'm not totally convinced that OAI or Anthropic have no moat, even as open models catch up to the frontier. Serving and inference are still hard problems when you're talking about a 2 trillion parameter model. Fine-tuning, if that remains a realistic need for businesses, is also a difficult infra problem at that scale. In the most bearish case, where there is no competitive advantage to using their models, big labs still have an advantage in this area.
Maybe there is some threshold where the price/quality math for your standard business tips in favor of smaller models and self-hosting the entire stack. I'd certainly love that.
8 replies →
A race to the bottom is where you lower standards, wages, or regulations to cut costs and attract business. What's actually happening is the opposite: a race to the top. Every model is trying to get better. Simultaneously they also happen to be getting more cost effective, but it's sort of a coincidence. Companies still require very good models, but they are not picking the "absolute best at any cost" anymore, because it turns out "any cost" isn't worth it.
Even if I pay for a true large model, I am going to prefer an open model hosted in Europe/my country, not them.
Maybe open AI and Anthropic could just license their models to run on your own hardware. So a fixed cost instead of per token pricing or subscription with limits
Fixed costs are already available via PTU reservation and afaik most serious enterprise projects are using this
This would require them to give you a copy of their models, which would quickly be leaked on the internet, and they know that, so it won't happen.
1 reply →
This is the future, because the market will demand it.
I think there will be separate and huge markets for models, hardware and compute. That will maximize competition and innovation.
Why? Because even Blind Freddy can see the huge usefulness and power of these (and future non-LLM) models and no-one in their right mind is interested in becoming OpenAI's or Anthropic's bitch. Those companies have tickets on themselves.
Given the recent behavior of tech companies and the US administration, no one trusts either anymore.
But think about the safety!! (/s)
They have good friends in the big ballroom to not allow you to use something cheaper and be locked in on them for your own safety
As a taxpayer, the only silver lining here is that we won't be paying for the ballroom... /s
The schadenfreude is that, at least for OpenAI, they were originally set up to make open models.
They were set up as a public benefit company and their returns were capped at 100x. They (well, Sam) went out of their way to put themselves in this death march to the IPO. If they had just done what Mark Zuckerberg did with Muse, they're not in this position.
Do you know how bad you have to be at the tech business to make Mark Zuckerberg look like a prudent-yet-visionary leader?
Yes, but also remember what happened to Mark trying to do a currency?
You're giving him too much credit if you think he figured anything on his own - this is the guy who thought 'metaverse' was a good idea.
3 replies →
I know a few Australian devs who work at places also moving, or already have, from Anthropic… and not because of cost but because of Trump’s edicts to ban non-nationals using AI.
Think about how crazy it is for non-US companies to use American AI providers - their marketing boasts that you can treat their models as co-workers, assign tasks, invite them to slack annd video calls, etc. Taken at face value, would you hire someone remote who lived in a country that commonly does random shit like deciding whether or not remote workers aren’t allowed to go to work?
[dead]
I know what happens in many big companies and not a single one is moving away from Anthropic/OpenAI/SpaceXAI.
> Unless they both dramatically slash prices then they’re in big trouble
False, they have already done so many times.
> Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any hope at a successful IPO.
False, margins are higher and I can have a formal bet that prices will go lower.
> However the cold reality for both is that there is zero moat to a model anymore
False, LLMs are not fungible and there exists a natural moat. I like the behaviour of Fable, not the behaviour of Opus - the fact that many people speak about this is evidence.
I don’t know what your sources are, but I work on AI at a Fortune 100 and open models now make up >90% of our internal token spend. Used to be 100% closed before this summer.
6 replies →
Not sure who you’re talking to but the NYTimes reporting clearly refutes your statements that nobody is doing it with clear facts. It’s been a tidal shift in attitudes over these last few months and the messaging back to OpenAI and Anthropic has been clear. Slash your prices by an order of magnitude or you’re done for most use cases.
We’re heading into corporate budget season for 2027 when all this is coming under a huge microscope in boardroom after boardroom across the country at a terrible time for companies trying to IPO.
> like the behaviour of Fable, not the behaviour of Opus - the fact that many people speak about this is evidence.
I'm wondering how much of that is the harness vs the model. Overall, the 'feel' of a model seems to be largely due to the harness than the model itself.
They kind of are fungible, up to a certain level of task. And much like most software developers don't need to exercise deep comp sci skills, most sw engineering doesn't have tasks that require the best models.
> I can have a formal bet that prices can go lower
Is this a typo? Have you actually made a formal bet on a prediction market or something to put your money where your mouth is, or are you just saying that you could? There's a lot of things I could plausibly make bets on, but that doesn't mean that they're likely to happen.
2 replies →
[flagged]
4 replies →
If companies are really doing this, then we're saying they have no problems spending tens of millions to get somewhat decent TPS and then having their employees complain they are timesliced and getting lots of timeouts because their org has 500 employees?
This is all spitballing, but I'd wager it's a third the type of workloads, a third hedging against your business depending on a single external provider, and a third trust.
Not everybody is coding or doing work that lends itself to burning tokens for warmth. Reuters for example seems to be more interested in using it for research, editing and formatting citations and the like. There's only so much of that work that needs doing, it doesn't always need to be real-time, and they probably don't see it scaling exponentially. They also need to be very aware and in control of their model's biases, or they risk it compromising their work output.
It's widely expected that all of the major providers will need to - and surely want to - drastically raise prices to justify the ludicrous amount of capital they're burning. Multiple companies have already talked about how their AI costs have exploded, and from what I understand that scale of enterprise is paying API rates. I would be disappointed if big business wasn't having a think about what that liability could look like. It's one thing to be reliant on a relatively "stable" vendor like Microsoft for Windows and Office, another to get AWS sticker shock, and then this is promising to be an order of magnitude worse.
Then just plain trust. What if ChatGPT starts recommending your competitors products, or the USA bars export of Anthropic's latest model (again, but for real this time), or they stop serving a model your business now depends on, and so on... That's a lot of risk to leave outside of your control.
Open weights != self hosted
At least accord to the MIT study last year, most employees are using their own AI subscriptions to do work.
Keep in mind that a vanishingly small number of workers are SWE's churning millions of tokens daily.
1 reply →
I'm not sure why someone hasn't developed a company offering services that distributes AI across all idle or under-utilized VM's and PC's for enterprises in order to serve open sourced models. Outside of the electricity bill, there's no additional expenditure and you get the AI.
We've all seen the office spaces where there's 200 empty computers on a floor. Combined, it's something like 500 cores at ~3 Ghz each and around 3 TB of RAM. The networking is already there and software like exo already exists.
4 replies →
There are most definitely is a moat - but it works both ways. The railguards in the models create moats keeping customers out. And the cost to build a modern agentic model is in the 10 figure range and growing. This is an expensive arms race that is going to create moats.
But most commodities are the same way. It’s super expensive to drill for oil. I need oil and I’m in no position to mine my own because of the massive capital investment. But it doesn’t stop it from being a pure commodity.
I couldn’t care less which company drilled for the oil… it’s all the same to me. Models are increasingly no different.
OpenAI and Anthropic are a gas station saying “buy our gas for 10x the price!” When the world is looking at them saying it’s just gas, we’ll take the cheaper brand. We’ve tested your gas and it’s really no better than the stuff that’s 1/10th the price.
Thats why their present business plan is screwed.
20 replies →
OK but even if F1 teams are a very expensive arms race it doesn't prevent me to bike to shop cheaply. You eed to have a moat around what people need.
1 reply →
> the cost to build a modern agentic model is in the 10 figure range and growing
Source? The proliferation of labs building competent models would seem to suggest the opposite.
1 reply →
The open models are good because of distillation, which the US labs are actively working against via not revealing CoT ever and now you can see with OpenAI Astra 6 not even having a lot of CoT equivalents being emitted as tokens. Once the anti-distillation stuff is in place the open distillation models will probably start having larger and larger gaps.
If the companies survive the next few years, which they probably will because they represent too much of US economic growth to allow them to fail, this gap will keep on expanding.
Starting from zero without distillation is a lot harder, a lot more expensive and a lot more work. OSS models is what a laggard does to get adoption. China's gov't might keep on sponsoring it as a counter GPU embargo thing, but when gov't get involved, usually the other side gets involved too.
As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.
If we accept the premise that the top Chinese labs are simply distilling and can't compete otherwise: why don't US labs simply do the same thing? Distill their own models and slash their costs by 99% while keeping the same quality output. It should be a piece of cake if even the open labs can figure it out, after all.
One way or another they're getting the same results as proprietary labs, with a fraction of the hardware for a fraction of the cost. OpenAI can't keep raising funding rounds of $100billions to subsidize their compute costs and get results by brute forcing parameter count. And if they're having trouble keeping up with Chinese labs' efficiency, maybe they should stop worrying about distilling and instead hire some of the smart people responsible.
8 replies →
If that was the case the major labs would have done that.
They’re burning cash like there’s no tomorrow. They desperately need to show that they have a real business and not just a giant burning pile of cash doing academically interesting things. If they could simply sell models that are 95% as good at 1/10th the price they’d do that. They’re losing the enterprise sector because they’ve not done that.
3 replies →
Basically your argument relies on two claims, both of which must be true.
1. Competitors to OpenAI and Anthropic are good because of distillation.
2. OpenAI and Anthropic will come up with some methods for preventing distillation in the future.
Both of these are dubious imo. For RLVR tasks like coding in particular, you definitely don’t need continuous distillation to improve, otherwise OpenAI and Anthropic themselves would not be able to improve because there is no better model to distill from.
Look up how many Chinese are pursuing a computer science degree.
Chinese AI labs do well because they have the best AI people coming out of a huge talent pool.
4 replies →
Yes, evidence is needed but especially for the claim that distillation is what makes these open models good. Serious citation needed.
Think about it: even if they distill the shit out of frontier models, the model still gotta learn, right?
If anything, as you can see from the K2 Horizon release, aggressive (self-proclaimed) reliance on distillation does not result in a model that has remotely any frontier capability. Try asking K2 Horizon to write iambic pentameter for instance, or even give it the car wash prompt. I tried both these on the Q8 quant for the 7B model and the results were depressing.
3 replies →
Do the western labs not train on the open weight models? Is it a one way distillation?
Open source does not apply to AI and we should discourage anyone from using that term. All the models are opaque and proprietary. You cannot go into any source code and fix bugs, or add features, or study it to learn more. It's not the same thing at all as open source software.
however, you can test the model, by asking it things. What are you concerned is hiding inside of that big block of numbers? Order 66?
this also applies to proprietary black box software. You can still test it, interact with it
True, openweight is the correct term.
I swear, Qwen 3.8 27B @ Q8 is smarter than Sonnet 5 most of the time. Why wouldn’t corporate America self host at this point, especially with better options like Deepseek Flash and GLM 5.3 flash that’s a middle ground between Sonnet and Opus
@q4 is definitely smarter than sonnet from what I’ve seen so far. It’s even caught problems in code made by fable, when using it as a code reviewer.
Agreed. And conversely, American models can also just as easily be secretly influenced for bad things, or be more tightly controlled by the government, to corporate America's own detriment.
America's own demise will be made in America, stamped by American laws
2 replies →
Post-IPO I'd trust the American models far far less than the Chinese models.
The most insidious advertising in the world is about to be surfaced as people use LLMs to look for product recommendations.
5 replies →
> Why wouldn’t corporate America self host at this point
Because they've been trained to think "cloud-first" for a decade?
Still a ton of non tech companies doing a digital / tech transformation out there too lol. Maybe some so far behind they still have the real estate and rack space to get ahead on this one
is this actually the case? I haven't kept up with the small models
but if there's roughly Sonnet 4.6 level capable open small models, then I'd be impressed
Qwen 3.8 27B is the real deal BUT remember to use froggeric template and/or medium reasoning.
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
Qwen 3.8 27B is better than Sonnet 4.6
Qwen 3.8 27b has the juice. Try it.
Why not Luna?
The financials are incredible. If you spend one engineer's salary on hardware, you get a system running local AI that can multiply the efforts of an entire (small) team of engineers. It's a very "you can't afford not to" situation. Even considering the hardware prices today.
> Some U.S. firms remain reluctant to use Chinese A.I. models because of concerns over regulation and data privacy. AT&T researches Chinese models but is not using them, Mr. Markus said. Instead, it is working with popular alternatives made by American companies such as the Gemma A.I. model from Google and the Llama A.I. models from Meta.
This makes sense since corporations require legal certainty, and using an open model from an American company (probably) provides them some level of indemnity, and also someone to sue.
The current U.S. regime is also replacing some amount of that legal certainty with regime fealty. Picking Chinese options over American ones probably runs a risk of upsetting their leader. I've got to imagine American companies are weighing this factor in their decisions.
That's certainly one factor, yes. But even before getting to that part I think the bigger issues for Big Corp legal teams is mostly around the legal ambiguity of the models themselves. What representations are made about the training data? What jurisdiction governs the license? If somebody alleges that the model infringes their IP, what rights does AT&T have?
Counterparty risk is a lot more straight-forward to evaluate when dealing entirely within the US, with US companies.
3 replies →
Regime fealty has always been there in US. The current admin is just more corrupt and throughly incompetent at hiding it.
Exactly. See TikTok trouble as example and quite honestly, try a local open source LLM and ask it to use profanity, paint nudes - the LLM doesn’t answer the question of it is from OpenAI or Google.
The thing is that needs more attention is reverse engineered a LLM which is highly fascinating. I tried it, but it seems I am not there yet to put it mildly. It requires serious effort.
I am just speculating but can LLMs be sleepers? You write software and it seeds traces here and there under certain conditions that pose a serious security risk.
Or a kill switch?
I don’t know. I distrust Chinese LLMs but even more due to training data.
It is after all not a Western model. Different biases and the might be subtle but nevertheless substantial.
In short: no open source LLM may be usable without additional Finetuning for certain valid use cases.
The real value is versioning and autonomy as well as lot more stable answering despite model rot.
Also testing and the supporting systems are easier to maintain.
It is mainly an infrastructure challenge.
Since most folks still somehow review the code, I find it highly improbable, it would become visible quickly. Unless say compiling Unix parts for example. But thats compiler work and not llm.
2 replies →
Read the license again.
The license you receive when you download Gemma off of Google's website is not necessarily the same license that AT&T gets when they deploy Gemma as a customer service bot. That's the whole point. AT&T can work directly with Google for a licensing and legal framework that provides certainty.
Spent most of today reworking a rack and rig of gpus for all of our internal ai work… our big server is 8 rtx 6000 pro and 3 psu, I definitely feel I made a mistake not upgrading our wall power to 240v but so far we have multiple 30amp 120v and with 3 PSU uninterrupted power we have been very stable. My big upgrade will be moving to epyc motherboard from threadripper so we get gen5x8 with bifurcation instead of what we are stuck with today gen5x4 due to bios limitations . What has been so encouraging though is first deepseek v4 flash at 200+ t/s for single user and much more in aggregate- now on qwen 3.8 flash next for image support and eyes on glm5.3 flash for some testing … next is realistically considering co location and quiet a sizable loan to scale this to real hardware instead of miner rip vibes
What kind of cost are you looking at and how many people could use it? I’d be interested to know what the payback period is like, because Claude code is getting ridiculously expensive.
Bought the rtx 6000 pros one per month starting in January- since they doubled in price I wish I just used a line of credit back in Jan to buy all of them. For me it’s about keeping internal company content internal. Slack channels etc with tools for teams to use. Triage tools for Zendesk etc.
> AT&T turned to artificial intelligence models from Anthropic and OpenAI in recent years to help with customer service, call transcription and coding. [...]
> By May, open models accounted for 20 percent of AT&T’s A.I. use. That has since risen to 40 percent and may jump to 60 percent in the coming months, Mr. Markus said in an interview.
This is missing a crucial detail. We know they "help with customer service, call transcription and coding", but which of those have been upgrade to open models?
Call transcription is trivial to do with open models. I can run Whisper or Parakeet on a low-spec laptop.
"Customer service" could mean a lot of things, but it sounds feasible for open models too.
"Coding" - they might go to open models for that, but I expect the costs involved in paying for closed models for software developers within AT&T are a fraction of the costs involved in transcribing all of their calls or handling aspects of custom service for millions of customers.
From later in the story:
> AT&T researches Chinese models but is not using them, Mr. Markus said. Instead, it is working with popular alternatives made by American companies such as the Gemma A.I. model from Google and the Llama A.I. models from Meta.
Gemma 4 is great, but really, Llama, in 2026?
> Gemma 4 is great, but really, Llama, in 2026?
I'd assume the author is just getting confused because of ollama and llama.cpp and all the other ecosystem "llama" that are still in use. Llama really did kick off the open models thing
I'd love to! For real coding though, SOTA models barely get the job done. It wasn't until Opus 4.5 that you could really get decent results.
I'm sure this will change (and I can't wait for it!) but as of today, open models might be fine for summarizing and writing docs, but you need SOTA to work on code if you want to be competitive.
Every time I check in on this I hear a more recent model is the one where they started doing good work. I'm excited to here that Astra is where it got capable enough to work on code next year
I think it truly just was Opus 4.5 where LLMs became usable for coding.
2 replies →
I do think there are now open weight models that are on par with (or beating) Opus 4.5 by now (e.g. Kimi K3, GLM5.3). But yeah obviously the frontier closed source models seem to have pulled away once again, so open weight seems to be a few months behind right now (which might be too long to wait for a lot of people!).
Those two you mentioned completely demolish opus 4.5. It's not even close. I'd say they are between opus 4.8 and opus 5. And better in some tasks.
1 reply →
You're not corporate America (and trust me, I mostly mean that as a plus).
I also work in software, and while I vaguely disagree that open models can't be used (they absolutely fit into productive niches here, and holy hell are the last generation [ex laguna s1, kimi k3, glm 5.3, etc] actually decent) - I will agree that SOTA are a better fit for software development, especially when used in conjunction with an already very expensive employee who's driving them.
But for "Corporate America"... no. You absolutely don't need SOTA. They're doing things like transcription, summarization, customer interaction, minor technical tasks like form creation in existing tools, report generation (ex - powerpoint, pdf, docs, etc) and other general "white collar tasks". Think about roles in business that are in the 60-85k compensation range.
It's mostly busy work that keeps existing processes flowing and the business on the rails. Important, but not research/novel.
And cheap ai... is a wonderful fit for a lot of this. No one wants to replace an employee making 80k with a less reliable AI that costs 45k a year in tokens (SOTA). But they're absolutely willing to drop 2-3k/year on AI (~100/month - right in the open model cost range) for that employee if they can get a 10% bump in productivity or happiness.
Greg Kroah-Hartman from the Linux kernel team said recently in a talk (https://www.youtube.com/watch?v=_MwMLPmMccs) running open models locally is already good enough for finding Linux kernel bugs and generating patches to fix them. He said that the false positive rate he's encountering was only one third to one quarter IIRC. He doesn't provide any details but it didn't sound like he was running some kind of huge server or something either, just a desktop workstation of some sort.
"real coding" carries a lot of the weight in that comment.
Seems like if you ask 5 different people what "real coding" means you might get 5 different answers.
Not everyone is building the next framework or compiler.
Self-hosted Qwen 3.8 @Q4 on my RTX 3090 can produce beautiful functional CRUD pages and apps all day long. And that is 90% of the "real coding" being done in corporate settings.
The quote in the article about Mazda vs. Maserati captures this. Many might want the Maserati and drool over its specs and capabilities, but balk at the cost and how often are they really going to run it up to full performance limits on their daily commute to their cubicle?
Yeah, I totally agree. I'm sure these open models are more than good enough for non-dev work. I'm also sure they'll be good enough for dev work soon enough (and some people are saying the latest already are). My point was that given the difficulty of writing great code and dealing with large systems, SOTA just recently emerged as a viable option. I expect open models to catch up soon.
Good luck to the poor shleps trying to make a living performing "white collar tasks" I guess, right? They can all go be poets or painters...
I generally use sonnet 5 for most coding tasks, a lot of coding is really pretty straightforward.
They couldn't pick a more sinister headline for such an awesome technological development.
Ah, so you're one of those hippies hooked on free software too, yeah? What's that your smoking there? Emacs, huh? what's your OS? Linux? I knew it.
Corporal, put him away.
Its almost as if NYT has had a hawkish agenda for the past few ...generations.
I definitely see a future race of who can accelerate open models and their usage locally with realistic desktop hardware.
I fully await my ai in a usb box. The models are plateauing and some clever company is secretly working on this already I’m sure of it.
Imagine baking in a model weight set in ROM with an analog computer doing what otherwise takes way too much power in digital form.
My long term guess: A&OAI will move away from being interference providers to just training models and then licencing the models for local use
> licencing the models for local use
Who's going to be the new Bill Gates, with a vision for "a GPU cluster in every home?"
Jensen Huang?
I'm not seeing it. Corporate America needs someone they can sue if anything goes sideways with AI given the rate of change and legal ambiguities. It took years/decades for actual, real, open source to be widely adopted in corporations for the same reasons.
Corporate America has been using open source for decades, and it wasn’t anywhere as slow as you portray. This argument simply doesn’t hold water.
Besides, in their present rather dire financial state there isn’t much to sue these companies for anyway cash wise. NYTimes is suing on IP grounds.
it holds all the water, to this very day using open source in my client's projects requires approval from legal.
1 reply →
The main indemnity that providers like GitHub Copilot provide is against being sued by the copyright holders whoes works were "transformed" to make the models. They generally don't offer that for the Chinese models, at least in my region. I can't see any reason other than back-room deals making it so: they're all using stolen works, and all could spit out memorized fragments. But if you're talking specifically about the fitness or otherwise of the actual software these models produce, that is generally handled with the use of sacrificial employees.
Not seeing it either. I do see a torrent of astroturf saying everyone is using self hosted LLMs, but in reality nobody I know is.
I had a conversation with a large group of friends and we independently came to the conclusion that Openai/Claude does not deliver more than a open source model. It takes about the same and the quality is about the same, and this does not mean it is good
I'm just glad the dweebs who were parroting "OMG, Distillation attacks!!1!" have been empirically proven wrong, and discussion around open-weight models are more rational now.
The thing is, for big companies (or even small ones owned by PE, which is MOST of them), it's not just about cost. The big thing is risk.
In my experience as a tech diligence assessor for PE firms for the last 7 years, investors really, really don't like companies being beholded to single entities that they don't control. Anthropic and OpenAI have demonstrated that they are not trustworthy, or predicatable, or finanically safe, or even capable of hitting three fucking nines. Investors know they need companies to be on the AI train, but they really don't like vendor lockin to the big AI companies. Every diligence I get asked "how easily can they change models?"
I think when open models reach 80% or 90% capability (or maybe even less!) a whole lot of companies are going to say "almost as good with way less risk is a better deal".
How does this compare to being beholden to a single cloud providers like AWS or Google Cloud?
Has there been a sea change in how investors view these things in general, or is it only AI?
I can’t personally tolerate the AWS/GCP interfaces and clickops.
But in terms of reliability - uptime, product, legal - they are in a different league.
There were exactly 0 instances waking up to a product decision at AWS completely breaking your product or workflows.
Well the big thing that stands out in the reliability front is that AWS and Google cloud are typically stable, certainly more than 3 nines. Meanwhile openai seems to roll dice on every request to see if they’re going to return a 500 or not.
IMO this will be a blip. There’s a lot of talk in the wake of all the Uber handwringing about token spend. Legacy enterprises want to look innovative to Wall Street without spooking them, so it’s easy to hop on the narrative and “show” that they’re innovating in a cost responsible manner.
This feels reminiscent of the big push to RAG a few years ago. And, more broadly the skunkworks projects that big companies tout in the press before they end up killing, when the operational overhead becomes too much for their liking.
Ultimately, the narrative is good for the consumer and the enterprise. It’ll mean OpenAI and anthropic will have to keep prices low. But ultimately, in the course of the next 10 years, I don’t see enterprises wanting to do this themselves. It’ll just be simpler (and eventually safer in their eyes) to send traffic to the big labs.
I don't see this being a blip, for multiple reasons.
A) the uptime matrix between Github, OpenAI, and Anthropic means that we've faced multiple entire days of not being able to ship org-wide due to our reliance on automated code review and other tooling. Every cloud service baked into our CI pipeline becomes a point of failure. We can't live without AI anymore, but it too often either directly or indirectly gets interferes with our ability to ship, and I don't see this improving any time soon.
B) Locally hosted AI has serious advantages with regard to PII/sensitive data management, and there's not much the frontier models can do to overcome this. There are so many things I want to build and let loose in a sensitive data environment but can't due to data governance around frontier models.
C) Anthropic and OpenAI cannot keep prices low forever. They're still burning insane amounts of cash and at some point, they're going to have to transition from growth mode to profit mode. They're already juicing their sales pipelines to the max with introductory pricing and other things to get people in the door. But those are all short-term online marketing plays.
Totally hear you on point a/b!
I think ultimately the folks with the purses won’t care enough about a for it to be taken seriously, even if it’s an engineering bottleneck.
B definitely has scope but still smaller than I’d expect. When I was an intern at yelp, I was migrating us off internal credit card management to Braintree/stripe. No one would have imagined outsourcing that in early 2000s. There’s a long tail of stuff that you can’t sent to a 3rd party; but for most use cases it’ll suffice.
For c, true; but this is expensive for everyone, including the Chinese model companies that are trying to undercut Open Ai / anthropic. In the nth degree, i think the field will bring the cost down to the place where it’s manageable (see the existence of the cheap Chinese models). The real question is the $$ spent on pushing the research forward at scale.
A lot of corporate use doesn’t require SOTA models. Things like summaries or writing comments in Slack are light work. For those reasons, local AI makes a lot of sense for some companies.
The article clearly distinguishes between open source and open weights models, and the headline and parts of the article state that it's open source models that are having a moment. But it doesn't list any examples. The only model families explicitly named are open weights only.
Could someone clarify whether there are actual open source models that are competitive with the likes of Gemma (mentioned in the article), or is the headline just wrong?
Anecdata:
I use opensource models at work because my work is too cheap to spring for a $20/mo account for me. Since HuggingFace models can be run on my laptop now (still very slow though), nothing is leaving the 'secure environment' and so I can actually get work done (instead of the 'old' version of coding and writing - google).
You still need the models to be able to perform web searches, don't you? In which case the data goes in and out of your machine and there is risk for prompt injection attacks. I think it's needed at least for documentation purposes.
It's just Linux all over again. Except nowadays open source isn't a "cancer", so it will happen faster.
Been wishing my company was doing this. We are locked in with chatgpt for chat and claude for coding. I have a 36gb m4 max laptop but I'm not allowed to run anything on it even though it would be totally free. Absolute pain.
Why do you care for company stuff? Local models would be slower and weaker.
Easy prediction: LLMs will get shrunk down further and further until GenAI is just something that ships on a chip as part of your hardware. In the future it will seem quaint that we needed a network connection to talk to our LLM.
Adoption of open-source models to my mind is a similar step in that direction. In all cases, the goal is to become untethered from a mercurial vendor.
Like taalas.com (very recently acquired by AMD), or cerebras.ai (whole wafer is a chip)? As you said, I also think that is one of the main direction many companies (and academia) is moving to.
We're gonna start baking in models like TTS with thousands of voices available in any language as a chip on device. They just need to hit 99% accuracy and then it's a done deal.
Yeah, I'm looking forward to this actually.
https://chatjimmy.ai/ blew my mind at how fast etched model weights can be.
For on-device LLMs, there's a point of diminishing returns, meaning you don't need to have the latest frontier model for most operations.
I currently have a small TTS model running in the background on my machine through which my agent(s) speak to me as they work. If that can be baked into an ASIC along with a few thousand voices in every major language then it should just be a utility chip on your mobo for anything that needs it. And yes, I too, am looking forward to it.
for every small GenAI model there will be larger model or cluster of models which are smarter than small model
Maybe the bubble doesn’t come for all of us maybe it comes for Anthropic and OpenAI.
The amount of downtime the big AI companies are showing doesn't help. Do you want your company's sales and customer service to shut down every time OpenAI or Anthropic goes down?
This is a prescient concern and it all came about due to the initial export controls placed on Fable on a Friday afternoon with zero warning. It was partly Anthropic’s fault though. Engaging in criti-hype by going on an advertising blitz explaining how their new model is some kind of genius super hacker didn’t help their case.
Nevertheless, it happened. The problem is that businesses demand stability. Building an application, service, or business process around an API that can be turned off whenever a government demands represents an unacceptable risk.
It’s not a mental exercise when it has already happened once before and likely will happen again. If you control the weights and invest in your own hardware to run them (or rent it from a cloud provider), that risk can be mitigated.
I think what they're really getting hooked on is the lowest cost provider.
Which makes the Muse 1.3 launch this week particularly interesting, although to get the low cost version you do need to agree to share data with Meta.
I actually think they’re hooked on models they can fine tune.
You can’t further train the closed models. The open models can be fine tuned for your company. Big companies fine tune models on all the internal systems and documentation, not just through .md files (you’d blow up the context trying it that way) but actual fine tuning of open weights models. A low tier but open weights model actually beats frontier models when you do this for a specific task.
I think the frontier providers need to have a way to isolate instances (bedrock style?) and allow fine tuning to compete. Big companies are absolutely fine tuning models right now and getting better results than even the best frontier models for their use cases.
Gotta love this article's use of both the terms OpenAI and open A.I. which are pronounce the same but mean very different things
What IDE/extensions do you use for open-source LLMs? I tried VSCode with ollama and lm studio, and the experience is very subpar to the built-in copilot. It's not very usable.
https://huggingface.co/Qwen/Qwen3.8-27B
No one I know uses the built in VSCode extensions anymore. It's all TUIs now. You can use Opencode as a TUI now for local.
Seconding the TUI usage. Opencode, codex, and pi are ones that I have / my friends have had success with using locally. I mostly use pi. It feels like the most "boring tool that just does its job" out of the big options.
I’ve been using Opencode, but recently migrated to oh-my-pi and have really enjoyed it.
I’ve also had luck with copilot-cli, but find that to be more limiting and it’s only really worth using if you are paying for GitHub Copilot (or your company pays for it, as is my case).
At our small company we are hooked on individual subs. But yeah a larger dev shop can't really pull that off and I get how they'd be dying by the token cost.
Contrary to what is mentioned in the thread there are no seats anymore, it is all raw token usage nowadays.
I have recently come to the conclusion that thinking for 2 seconds and using a cheap model with a slightly more detailed prompt works just as well as zero-shotting an idea with a fancy model. I work in science, and instead of asking the model “write a topic extraction algorithm”, I just say “hey look at this matrix factorization script I found in a repo, now make it use plotly and duckdb”. Have others come to the same conclusion here?
It makes me skeptical that the flagship companies are sustainable. Every company is going to maximize “fuel efficiency” to save time and money.
Then again, maybe the cheaper models have more markup for them, in which case they are probably happy w this arrangement. I’d be curious to know how the money making varies by model.
Really? It's worth it to you to spend 10 minutes thinking about how to prompt a dumber model to save $0.05? (not that open source models are dumber any more)
Its not just about saving $0.05. There are many legitimate reasons to not want the AI to do the thinking (aka architecting or planning) for you. In those cases where one-shotting is undesirable or unnecessary, closed frontier AIaaS has no advantages over open weights.
Yes? The difference is often 10X or 100X with very little time lost. I’m learning how to give it enough info pretty quickly.
Edit: I also have to read the methods anyway for scientific accountability/integrity anyways, so I may as well play that role at the outset.
Well its def not $0.05, I just started using claude sonnet 5 and Ive found most simple questions might be 0.05 cents, a unit test is something like 0.10 -> 0.20 and small features and classes get into the individual dollars. Sure its a lot faster but at the end of the day its not cheap.
Plus there is something to say about being in the drivers seat, youll have a much better idea of how it works instead of needing to talk to claude and hope its correct. Since most LLMs also not very good at ideas even in my experience with better models its better to think for 10mins, youll get a much high quality result
1 reply →
I think using open-source AI is no longer about API cost but about company survival.
Take Anthropic for an example. Anthropic has successfully destroyed customer trust, at least for me. DHH in a recent interview mentioned that Claude refused to translate an article about immigration. Not summarize. Not editorialize. Translate! I think this reveals an unacceptable level of paternalism: Anthropic fundamentally believes that it possesses a moral authority superior to the people actually paying for the API. If such basic and mechanical translation is already too sensitive to touch, the goalposts have moved from safety into outright censorship. What prevents them from quietly deciding tomorrow that your proprietary business logic, financial data, or legal documents cross their invisible moral line?
Let alone how Anthropic treats Cursor and Figma - not that they are wrong as companies are free to compete legally, but nonetheless it shows that companies can't outsource their intelligence to a potential competitor.
I get what you're saying and it's concerning how much power these big labs have amassed and how little transparency there is in what they do with it...
But I doubt this a major factor in the trend. I just don't think it's something most corporate users run into. My understanding is these guardrails are negotiable for enterprise customers anyway.
And, not for nothing, but if I owned a human-powered translation company I would've refused to translate it too.
I like the Claude constitution overall - I hope it becomes something representatives vote on and amend, to avoid the centralized corporate censorship you describe. In the meantime, I am fine with it abstaining from doing DHH’s bidding, especially because there are so many AI alternatives.
Ah, yes, I'm sure the article that moral paragon DHH wished to translate was not at all harmful, and that this was a good-faith effort on his part /s
While I agree that Claude can be overly paternalistic at times, how should it respond to a request to translate, say, bomb-making instructions? It's reasonable to me that it might refuse this.
How do you know what was in the article?
Also, you see zero distinction between hearing opinions on political topics you might find objectionable, and building a bomb to kill people?
1 reply →
The innovative edge markup already faded and the race is to the bottom, more features, more reach, less cost. It's going to be extremely hard to recoup those giant investments. No, the bubble won't pop, it already popped and morphed at the speed of AI that we didn't even notice, money just realigned, llms keep pushing the frontier, and peripherals are gaining momentum
The race is still on
How is the NYT's copyright lawsuit against OpenAI going and why have you abandoned your start witness Suchir Balaji?
Have you been brought into line? Open source AI also violates copyrights.
Oh no! I'll make sure to tell the Chinese companies about the copyright risks.
Most open source fans are also hostile to copyrights existence and are openly IP abolitionists. As such, they collectively respond with "good."
This is actual communism, and the fact that Bernie Sanders and every other member of the DSA isn't actively fighting for open source and is often fighting against all AI shows how fake their purported movements are and have always been.
If IP were entirely abolished tomorrow with no other change to our economic system, you would still not have universal Healthcare (maybe drug prices would be lower, though), Elon Musk could still donate however much he wants to get his preferred politicians elected, fossil fuels would still be used in amounts that destroy the world, etc. Very importantly, AI would still be used to try to manipulate and control the public.
The fact they prioritize other fights more than OSS, and have a rather dim view of AI, is hardly proof that they are fake.
> Most open source fans are also hostile to copyrights existence and are openly IP abolitionists.
Open source licenses are only enforceable because of copyright law. How are you going to enforce GPL3 when you have no legal authority to say what people are allowed to do with your code?
There's no reason that one's feelings about AI can't supersede their feelings about open source. We all live with a complex tapestry of values.
While I don’t disagree that they should be fighting for open source AI, Sanders is hardly “fighting against all AI”.
https://jacobin.com/2026/07/ai-nationalization-sanders-liber...
1 reply →
Who cares about power user fans? They don't own the copyright.
Open source authors have always been protective of their copyright. There are numerous examples when drivers have been copied between BSD/Linux (I forget which direction) which led to huge flame wars.
The whole point of the GPL is that it uses copyright and copyright assignment to the FSF to protect what it calls software freedom.
BSD authors are very upset if the attribution clause isn't observed. And so on.
It is communism to exploit poor open source authors? I have to read Marx again.
Wanting to replace monopolies with competitive markets is the strangest definition of "actual communism" I have ever heard.
Anti-capitalism does not automatically qualify as communism. Notably, if IP were to be abolished, it doesn't belong to anyone. The expectation that open anything includes some sort of DRM-like content gating is counterproductive.
It's possible that LLMs will be a commodity in the future. Just like airline industry, AI will be tremendously important for society, but AI companies will not be making lots of money. It will be Nvidia, Micron, Dell and others shovel makers making money.
My hot take is that open models don't really save you money and introduce more router complexity and security risk (because you're now sending your company data through more less trustworthy providers). Look at cost per task not cost per token and the pareto curve is largely owned by closed models.
Just use Fable 5.1/Opus max for the hardest problems, GPT Sol high as your workhorse, and maybe terra for async batch stuff you don't really care about. Gemini 3.8 High also looks pretty good and is quite fast if you're already a GCP shop. You can basically benefit from open models without using them because they force the frontier models to be cheaper.
We run Gemini fast for random end user queries for general staff.
We have our own on-premise inference server (quad MI300A) that runs Kimi 2.8 extremely well and we transitioned all heavy work to it since it's basically instantaneous for the whole team. It's a good enough solution and we will hit break even before the end of the year already.
Not everyone needs frontier models and availability is frequently much more important than a lot of companies realize.
I don't think the "don't really save you money" hot take holds water in every case.
Coding, maybe.
But for operationalized/repeatable tasks it definitely does.
For example I have a workflow that I was running in April that effectively would cost $30k in token spend for each full run.
However now, with GLM 5.3-flash, we've brought the cost down to $7k-9k with our evals showing we've had no loss in recall, precision etc..
Just use OpenRouter
This is not a viable strategy. You're effectively paying for hardware and/or cloud. Why do this when OpenAI and Anthropic are both significantly subsidizing costs to win the market?
Talk to “AI” executives at large firms and 95% of them are clueless sales types that crawled their way to the top. Then again, it is basically a repeat of IBM, Microsoft, Oracle, etc. Same dumb executives making decision to not get fired and enjoy their place at corp.
The "AI" executives at large firms are mostly product folks with engineering backgrounds, not sales, which should make us even more worried.
Some of the most insidious parts of AI infrastructure includes the embedding model. Corporations have already spent an outstanding amount of time and money creating embedding vectors that are closed source and not reproducible. This means that all their data is locked into whatever embedding model they chose initially.
I highly recommend utilizing an open sourced embedding model instead of paying for a closed source one. It's vastly more reasonable to run an open sourced embedding model as a first step. They're much, much smaller and, due to the overhead of network latency, and running it locally has almost the same speed as through an API even on slow computers.
I would even go so far as to say that closed source embedding models have a high risk of data hostage. If a team doesn't have access to the embedding model, the embeddings become useless. A corporation like OpenAI could, say, hike the prices to that model by 1000x and everyone would have to pay up or forfeit any utility of the data.
I envision a future where open source embedding models are shipped with relevant technologies and implemented by currently under-utilized chips like NPU's. A startup developing cheap microprocessors that can run them is an idea I would pay cash for. Or perhaps they will be bundled with security tokens.
While it might be impractical for all corporate teams to run language models, it is very realistic for everyone to operate an open sourced embedding model, at least in their private cloud. Better yet, utilize transfer learning on an open sourced one to train your own, that way the embedding vector is more secure against competitors and trade secrets.
What are people even using embeddings for these days? It certainly seems like giving an agent grep covers most of the use cases. Dare I say: grep is all you need.
how much of performance comes from inference time tricks like scaling, topn ect . maybe models providers are also in position to run their models vs running os models by a generic providerc
Google can make themselves the heroes of the AI story by releasing a 120B dense Gemma model.
They already released the transformers paper, and I'm sure they are now scratching their head about it.
Meanwhile the rest of the world tries to un-hook itself from corporate America. Too many problems coming from the USA lately - it is not worth it to help sustain this anylonger. Canadians have realised this - others are realising this as well right now. Mr. Trump "no more forever wars", starting another forever war.
[dead]
[dead]
Using models that are created as a psy-op by foreign countries will end badly.
Using cloud hosted models that are created as a psy-op by domestic billionaires and crypto-fascists will end badly.
At least with self-hosted models the people running them get full control and don’t need to worry about the model or guardrails changing under their feet.