← Back to context

Comment by Gareth321

19 hours ago

> However the cold reality for both is that there is zero moat to a model anymore.

The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability.

On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision.

On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places. In California or Europe this could be $30-60k per year.

As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less.

The last major moat is the ability for subscriptions to scale with need. This means easily adding and removing licenses. This is far easier than purchasing extremely expensive hardware (and managing it), and selling it if/when internal demand changes. It's the same reason companies use contractors. The ramp up/down costs are very high.

The only real moat that local LLMs have right now is privacy.

I think you went from one extreme to another.

OpenAI and Anthropic rent their compute from AWS & friends. When we say large enterprises are moving to open weight models it means they are cutting out the middleman and renting the compute directly from AWS instead of giving OpenAI and Anthropic a margin.

> they have at least a 3-6 month head start

This is a moat of nothing. Our company still hasn’t gotten access to Fable so switching to open weight models would mean getting access to similar quality models. In some orgs they are still on 2025 models.

"Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision."

Depreciation is designed to _encourage_ purchasing of useful local tools, by incrementally matching fractions of the cost of the tool to the revenue it generates over its useful life. The fact that a graphics card might have a book value of $0 after five years of depreciation is a feature, not a bug.

Since the invention of corporation tax it has also had the benefit of offsetting tax over the same period, instead of just one big offset in the first year.

  • With all due respect, it sounds like you believe that when an asset is depreciated, it means the company gets to claim back the pro rata capex amount in tax. That's not remotely how it works. It's a business cost. Revenue minus costs equals profit, and that is taxable. Depreciation allows the business to declare lower profit (and thus pay less tax), but note that *the business is making less profit.* That's bad.

    I'm not challenging the concept of depreciation. It's a necessary tax function. I'm explaining the business case for local LLMs is poor.

Hmm yes but the equipment cost moat is artificial. This scarcity was created by the big AIs by buying up all the future production capacity. That works for a while but it won't last forever.

It's the same with the subscriptions. Local models can't compete because they're simply giving too much value for money. They're effectively subsidised by Big AI. Again something that won't last.

You don't need to self host to get the benefits of an open model. There are many hosted providers cheaper than OpenAI or Anthropic who can give you a SLA, ZDR, BAA and all the other three letter acronyms your compliance department needs.

The important part is if they break the contract or raise their prices you can always move to a different provider. You get lower cost and lower risk at the same time which is extremely rare in business. That's just not possible for closed models where your only options are the official branded API or Azure/Bedrock.

Your annualized energy cost estimates are off by an order of magnitude. 1kwH @ $0.1 (Texas) is $2.40/day if 100% utilized 24x7, California is roughly twice that per my understanding.

Privacy is non-negotiable for corporate. Even without considering costs or country of origin, we've seen from OpenAI that claims of AI safety are worth less than the (virtual) paper they're printed on.

All it takes is one incident, and all your company's internal data will start showing up in public users' chats. You can rely on a contract to prevent this, or you can guarantee it by using a locally hosted model you fully control.

When combined with the cost savings and good enough performance mentioned in the article, this can become a huge selling point.

  • > Privacy is non-negotiable for corporate.

    Corporate doesn't care at all about privacy. It's why the run outlook and windows and let Microsoft scoop up all of their company secrets. It's why they hand every scrap of data they have on their customers to salesforce and surrender their data to Atlassian and use Confluence and JIRA over countless alternatives.

    All companies care about is that when data breaches happen publicly they can point the finger at someone else.

    • Ah, Jira, the urinal of corporate cyberspace. You would literally be better off divining with some bunches of reeds than trying to figure out what the hell is going on with some stupid upside down Austrialian trash software.

      1 reply →

  • Is this not true for everything then? All cloud services, all Internet providers, everything within M365, OneDrive and Databricks?

    • Claude is different from S3. AWS doesn’t need to rifle through your files to stay ahead of the competition or to mine them for business ideas because the core business is overvalued and rapidly commoditizing. AI labs, on the other hand, have an incentive to exploit every last drop of data they can lay their ethically challenged hands on. And they’ve already demonstrated that they will do this, even when it involves blatant fiduciary violations (see eg, Anthropic / Figma board member scandal).

    • I have never heard of a cloud service of ISP actively taking corporate data and shoving it into other users inbox. If my company has a proprietary solution to a math problem that nobody else knows about but the AI service that we have been working with has access to it and Billy Main Competitor ALSO has that service and asks "how would one solve this math problem" the second the AI spills the beans, the suits in the c-suit would lose it and sue the AI.

      If however, our internal IT group didn't read the manual correctly and set up our cloud service access parameters incorrectly so that anyone who was trying to probe our defenses could access the formula, Jim from IT is looking for a new job.

      See the difference?

      2 replies →

You’re missing the point that 98+% of the use cases for AI don’t require the latest greatest model and are far better positioned to use the fast-follow distilled cheap models.

OpenAI and Anthropic are fighting to win a race (build the biggest baddest model) that has no prize. The prize is mass adoption at scale at the best price, which is why companies are rapidly shifting to open model. They don’t need to pay 10x for a model that’s provides no practical additional benefit.

  • Exactly, unless they solve reliability and jaggedness somehow and keep a moat with it, I don't see sudden brilliance with compunding errors leading to substantially more adoption. Most things that need doing in corporate america are quite simple but need reliable follow-through

  • Yea, they're really hoping that the 2% will help make up an outsized share of the revenue and that brand recognition will keep them going with the plebs.

> As for intelligence, the frontier models from OpenAI and Anthropic are still superior

I'll grant they are superior at least right now. But also, they are too expensive.

We ($work) are finding that it is best to build engineering discipline around AI usage (who would've thought!) and use the cheaper models like Cursor Composer.

Using Opus we can blow through an entire month budget in an afternoon, so while more powerful, it is no longer practical except for rare very complex tasks.

Re this point

> As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less.

I would argue that the reason they cost so little is because anyone can run open models and offer them as a service, so there's actual competition and the price is closer to cost. i.e. if the open models were just as intelligent as frontier models but cost the same to run as they do right now, the price wouldn't be higher (unless demand went up so high that marginal cost to provide more of the service went up, due to scarcity of hardware and or electricicy).

On the other hand, if what you're saying is the frontier labs have some pricing power due to their models being better, and that is the reason they are able to charge more than the companies providing open models as a service, then I would agree.

Actually, the moat is regulatory. Expect these companies to behave themselves in progressively more grotesque and sycophantic ways to get the federal government to make open/foreign models (and their output) illegal. After all, their very survival depends on it.

>On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription.

One of the basic questions/concerns here though is that it's not like the AI places are getting the GPUs for 10x less. It's true they have some economies of scale, but they also have some waste, and frankly in this particular case it's not clear they get that much gain over what a lot of businesses could achieve. The biggest traditional gain for central providers is that a lot of typical computing usage is burst-y, and in turn local kit might be underutilized. But with LLMs heavy users tend to use them all the time assuming their tokens allow it (and in the case of local hardware there's nothing stopping you, quite the contrary), they can use it directly interactively or leave them to go overnight on something too.

So it's reasonable to suspect that the reason subscriptions are only a fraction of the cost is that we're in a bubble seeing these companies losing money in an attempt to gain some sort of durable advantage. Just as every previous time, there is the chance that the music stops at some point, and they need to crank up pricing or pull other schemes to actually make money. Of course, it can be a good deal in the mean time, you basically get to suck down investor money for nothing, but it's also not unreasonable to at least be consider fallbacks. Even beyond questions of control and risk etc. I know at least a few places that are now genuinely considering questions like "what happens if a datacenter we depend on gets droned" that would have never had an iota of thought devoted to them even 5 years ago.

>On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places.

I don't think that's "surprising" at all, everyone knows about power use. And this seems like it gets heavily into what you're defining as "usable" and is also more useful to define in terms of cost-per-employee vs total. Obviously a bigger business will have a higher line number total even if the cost per employee is identical, but simultaneously can be expected to be making more revenue to pay for it.

If we're defining an average of a dedicated 5090 pulling 1 kW for every single employee (presumably some people wouldn't use it all the time, but others would then pull the compute for other work), running 24/7 (to cover people running stuff when they're away), then that'd be 8760 kWh per year. At my not particularly cheap New England location that'd be about $1900 per employee per year at the generalized residential rate (~$0.22/kWh), or $156 per month. That doesn't seem radical if it really does boost productivity. However, there is a lot of room to go lower. I'd expect a business to run backup anyway, and these days there are a lot of incentives to do that at least partially with batteries. That also opens up rate shifting as another way to pay back the cost. If we change to time of day pricing, that's 8 hours of peak pricing with the rest off-peak. 8 kWh of battery can now be had for a few thousand. And the off-peak rate is only ~$0.14/kWh, cutting the cost per year by about $700 to $1200 per employee per year. Solar power is also usually far more valuable to use yourself then sell back to the grid, and also continues to plummet in price.

None of this is to say that it makes sense for every place at all, but it's close enough to the the line that the math is at least worth exploring, or could at least lower the cost enough to be worth it given other things. It really comes down to how much extra value the company (or individual) expects to come out of it per month.

>In California or Europe this could be $30-60k per year.

Dunno about Europe, but at the kinda prices I see for California I'm really surprised more places aren't trying to move a lot of usage to battery+renewable.

>The only real moat that local LLMs have right now is privacy.

I don't think resiliency and control are things that can be taken for granted anymore, particularly on the global scale. War and terrorism is getting worse again. International relations are getting nastier, and governments have the power to just order places cut off. If LLMs aren't particularly valuable to a business, then why an expensive subscription? But if they are particularly valuable, then insurance is something leadership should be contemplating.