Comment by geophile

1 day ago

The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

- PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to.

- PC office productivity software destroyed expensive professional products.

- Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge market share from the mainframe world.

Ignoring the huge Chinese open-weight models for a moment:

- The training costs and resource requirements for frontier models are unsustainable. The high price, and social pushback, mean that the American companies producing these models are precarious.

- There are enormous financial incentives for research results allowing for cheaper, less resource-intensive models of high quality.

- Local LLMs on consumer hardware are akin to the PC hobbyist world of the 70s and 80s.

Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

Getting back to the Chinese models: They allow for new competition against Anthropic and OpenAI, basically SaaS renting out these very capable AIs much cheaper. That will just accelerate trends.

The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.

With open source projects, the benefit was that each individual could improve the complex system (e.g. Linux Kernel) interpedently, and over time the benefits accumulated. With models right now, there is just no way to do distributed training, or really, any large scale parallel way to improve them.

So whatever the short term strategy driving publicizing the model weights (e.g. potentially, to create a price war in order to put pressure on western companies and deprive them of the money they need), we can't ignore the fact that incentives and decisions could easily change in the future, and unless there is a way to truly decentralize models improvements - the party could stop at any time.

  • Nope, that's not true at all in China. Everybody in China is using AI, even the elderlies and the kids. DeepSeek is already earning money.

  • But different huge companies have different incentives. It is very much in Nvidia’s interest to have me running a powerful open source model on a $4k machine that they sell me.

    • Is it? When they could be having you running an even more powerful model on a $50k machine they sell by the pallet-load to enterprise consumers? We already see RAM manufacturers abandoning the low-end market in favor of server support. It's not clear to me that Nvidia sees personal GPUs as their best long term investment compared to selling millions of server-farm class machines

      4 replies →

  • Who funds the majority of cutting edge scientific research?

    Is it companies or is it governments?

    If governments around the world see LLMs built from public knowledge as pre-competitive as the public knowledge itself, then why wouldn't they sustainbly fund it?

    • Because the business model for governments is difficult. Income are mostly taxes, so you have to be able to explain constantly why it is a good idea to keep on spending budget on that funding.

  • There's a lot of individual effort of improving the models. See how many finetuned models and LoRAs are there on Hugging Face.

    • fine-tuning a model is very different from training the whole model in terms of resource requirements.

  • Soo? If they need a better version, the world can pool resources together, form a company that trains the model, then the company goes under and the model becomes open source again.

  • > The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.

    Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."

    There is more to society than capitalism.

    • > Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."

      > There is more to society than capitalism.

      I don't read GP like that. I read it as "we should recognize a situation of unstable incentives for an important outcome, and start thinking about other solutions."

  • Well that's kind of the point of the article. That in order to "win", the US needs an incentive structure that encourages open models.

    I'm not sure what that looks like though.

  • Why do people always bring up state support when it comes to China? As if the U.S. doesn't provide massive tax breaks and explicit funding to industry?

    It's on every tech post about China, as if it gives them some sort of "unfair" advantage.

  • Don't have a choice, will probably have to go open-weights models as currently, "AI" is gated using 'whatwg cartel' web engines.

    In the light of this, I am mechanically a proponent of very good open weights models, which I can download (for instance on on bittorrent) and run, slowly (the price), on local hardware.

    That would be for coding.

    If china puts its AI models on the same ground than US capital investment funds and big tech financial support (aka Big Tech international finance), they will very probably lose everything (know how, ML and inference infrastructures).

  • I am also confused by this point. The American government could force OpenAI and Anthropic to open their models, but then they would instantly evaporate, right? It doesn't seem like a choice that they can make, so framing it as a "winning" strategy doesn't make any sense to me. In what world could those companies have existed and opened their models?

    • They could, but they don’t have to. The Chinese have beaten them to it, and the rest of the world will benefit from it and the circle will be complete once the models get a little bit faster/smaller and the localized hardware does the same and it will, it is inevitable.

      The one thing that is sort of ironic or bad is that between Russia and the Ukraine there’s a large number of mathematically inclined people that if it wasn’t for the Putin war, their brain power working on AI models would have probably pushed open source down the road, even faster…

      5 replies →

  • In Russian opposition's mostly liberal discussions their school of thought connects several things together (sorry for not going directly to Marx's "General Intellect" and "Fragment on Machines" and using AI summaries instead ) - general idea of communism in China vs. techno-libertarianism of Thiel, Musk and the likes, and the Marx's thinking like:

    "Fragment on Machines":

    "he explores how human knowledge and collective intellect become embedded into machines, divorcing the worker from their own creativity."

    "General Intellect":

    "These texts are widely discussed for his concept of the General Intellect—the idea that society's shared, collective knowledge increasingly drives production rather than raw manual labor, and that this knowledge is alienated from workers and used as an instrument of capital."

    (note: my point isn't to pass any political judgement here, like what real communism in China or not real, is it good or bad, i just find it interesting that pure political discussions by people with no technical credentials bring AI as a major factor today)

  • > The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.

    Right... and there are two problems with this:

    1. Eventually the capabilities of closed-weight models will just vastly outstrip open-weight models if the underlying assumptions about compute and scale needed are mostly on the mark. So you can release open-weight models and they will have great use cases and applications, but ultimately similar to how you don't use an open-source phone or a budget Android phone from Wal-Mart and you buy an iPhone instead, you will see that although they "do the same thing" one product is clearly superior and you just have to pay for it. For this to not be true...

    2. then it incentivizes most (all?) companies, American, Chinese, or European to halt development of models because if you spend all the CAPEX and it can just be copied and turned open-source nobody will invest in that. Given that China is not halting development of proprietary models I believe the current strategy and the subsequent approach to release open-weight models is at best a stall tactic, and at worse a sign of desperation.

    Open source and the support and development models around it have been great. But folks are a little too dogmatic about it. Open-source software isn't a moral good, and closed-source software isn't a moral wrong either.

    • Imagine there’s a school where all the kids there are being tutored by the best. Also imagine a bunch of neighboring schools drastically falling behind that would need insane amounts of money to keep up.

      This becomes a problem because all the kids from the rich school will dominate the order schools. They’ll get even more money as time goes on from their kids paying it forward to the point where all other kids are bound to work for them.

      Now let’s say one other school does have the money for best tutors, BUT they know they’ll run out pretty quickly. Instead of trying to compete in a losing game, they decide to give every school in the world access to their elite lesson plan. Now, for a time, everyone will be on close to a level playing field. If the other schools improve upon their own lesson plans and keep sharing them with others, one day the elite school will wake up to find they are no longer on top. The parents have started to move their kids to other schools because the rich school is no longer attractive at the high cost they charge students

      7 replies →

    • Software being open source has many strong positive externalities. It advances human knowledge and freedom. If you don't think that counts as a moral good then I'm baffled by what you think a moral good is.

      1 reply →

    • Open source is in the tradition of humans sharing past knowledge, long-term we just can’t keep a secret it’s a time, honored tradition…

    • It's important also that open-weight isn't open source. If you can't download the training data (fully labeled), source code of the NN, and follow the README to build and train it yourself assuming oyu had the hardware then it's not open source.

      tldr there's no "source" in open weight models therefore they are not open source.

      1 reply →

Personally, I don’t think the general-purpose LLM as a standalone tool is long for this world, at least not in consumer-facing applications. I think when the economics make more sense, product designers will make things that people actually want to use that will pretty transparently handle whatever model interactions are necessary, when it makes sense. As a consumer, the last things I want in an interface are to a) be sycophantic enough to lessen my judgment, and b) be obstinate, obtuse, or argumentative, or generally just be something that I have to explain things to. I think a lot of tech folks are far more biased than they realize by the “ooh, neato” factor when imagining how nontechnical people might want to use things. And the weight of these tools just feels wrong for what a lot of people use them for: the thing that plays whatever music I feel like hearing absolutely does not need to be able to generate a volumes of fanfic about the movie that song was in. It’s abstractly impressive that something could do that, but it’s just not useful.

  • >As a consumer, the last things I want in an interface are to a) be sycophantic enough to lessen my judgment

    This is EXACTLY what people like/are addicted to about chatbots.

    My sister-in-law bombed an interview and asked AI about her answers to the interviewer's questions, chatgpt or whatever it was told her that her answers weren't bad, but that the interviewer could not see the gold in her responses. She said she felt much better.

    I see this effect with all the non-tech people in my life

    • I prefix many of my LLM chat sessions with this line

      > Chat rules : no sycophancy or over-agreeableness

      (But even with that rule it's still necessary to be discerning about the responses you get and to push back against points made, or words used)

  • > I don’t think the general-purpose LLM as a standalone tool is long for this world, at least not in consumer-facing applications.

    I use AI chat every day, I find it endlessly useful. It’s replaced google search.

    • > I use AI chat every day, I find it endlessly useful. It’s replaced google search.

      Extremely subsidized agentic search is very superior to Google at the moment, and of course it is. Google is a public company. The AI summary model has to work instantly, is likely as dumb as a 8T param model, and gives you incorrect details constantly. This sucks so much for Google. If you click on "AI Mode," suddenly the facts become more accurate.

      Of course, if I want a real answer I happen to go to claude.ai, set it to a the best model, wait for a minute, and use many watts of energy. Slow agentic search that takes many seconds, and is greatly subsidized, is certainly better. This should not be a surprise, should it?

      I think it was on a sub like r/singularity that I saw a post along the lines of "of course most people think that 'AI' sucks, as normies are interacting with 8T param models."

      tone: genuinely confused about the world, not criticizing

      3 replies →

    • I've found LLMs useful for surfacing popular recommendations. I also get the overwhelming feeling that it's all very early days still when the machine mixes together whatever was crawled into the weights with a web search or two and dumps it into a markdown blurb.

      I totally agree with the above that a more polished and less obvious use of LLMs integrated back into search engines may be more useful, but will definitely be more usable.

"I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now."

I don't think it'll take 10-15 years. Gemma 4 31B in the 4-bit QAT is competitive with the frontier of less than three years ago and runs on any high-end 32GB gaming PC GPU or a large-ish Mac.

The question is whether the frontier will continue to get better at a rate that allows it to stay ahead of the two curves of availability of consumer hardware big enough to run somewhat larger models and the capability of small models to compete with large ones. When the bottom falls out and GPUs/RAM becomes affordable again, the size of what normal people have on their desk will trend quite a bit larger than today.

I think there's a future not too far from now, where a 120B model with really good reasoning and a large context, but limited knowledge (necessitated by being small, you can't fit the world's knowledge in 100 gigabytes), can substitute for a frontier model on almost any task, just by giving it access to web search and documentation for the thing you're trying to do. A 256GB unified memory machine with sufficient memory bandwidth would comfortably run that 120B model.

  • I think the question is even a bit more nuanced than that. Even if frontier models can maintain a big gap that gap has to actually _matter_. If a local model satisfies my everyday use cases adequately then I may not really care that a frontier model is 5, 10, 100x better at ultra high order reasoning tasks.

    I think that reality is probably not all that far off for a huge swath of use cases.

  • Hell, Bonsai Labs 27B parameter model can run on phones with their ternary implementation which is quite efficient. Scale that up to frontier model parameters and it's quite likely we can run them on current laptops.

  • Came here to say that, my bet is that in 3-4 years you'll be able to run Fable-level of intelligence models on your laptop or maybe even on you phone

    • But isn't there the raw intelligence of a smart model and then the practical intelligence fuelled by how many parameters it has? You probably will barely be able to fit a 70 billion parameter model on a phone in 3-4 years let alone a 2+ trillion parameter model... so it depends on what you call intelligence

      1 reply →

Another important thing that made software usage and education available for most of the world was piracy. I remember as a kid growing up in a developing country, any software (windows, office, Visual Basic, flash, dreamweaver, etc.) was less than 1$. That allowed me to try out and learn so many things on my own without paying a huge amount of money for the license. And I think this is true for most of the software developers of my generation who grew up in developing countries

  • > windows, office, Visual Basic

    This is all (Microsoft) junk and so I wonder if you actually benefitted from this 'piracy'. And of course, it's well known that MS turned a blind eye to such 'piracy' in lesser developed countries, as they knew that they were gaining a future paying customer base.

    • I’m talking about late 90s/ early 2000s. Back then Microsoft was the state of the art when it comes to PCs. I remember using MS Frontpage to build websites back then and VBasic to build simple programs with UI and using MS Acess as a database. Of course I did benefit from it.

    • You're saying that using the most used operating system from the 90's, the 00's and the 2010's, as well as the number one producitivty desktop publishing software and visual basic would not benefit a user, especially when they had a near zero cost to understand the core concepts of end user computing?

      Really?

> I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now

10-15 years? The current rate is closer to 10-15 months.

15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.

Today, you can easily run Qwen 3.6 27B on a variety of consumer hardware. It scores 37 on that index.

Here are a number of open weights models that you can run locally compared with the frontier class models from 7 to 15 months ago: https://artificialanalysis.ai/?models=o3%2Co3-pro%2Cclaude-4...

I've run all of these models on my laptop (Strix Halo, 128 GiB of unified RAM); the bigger ones, like MiniMax M2.7 and DeepSeek V4 Flash, need to be done at fairly aggressive quants that will certainly lose some performance and not quite hit the performance of the unquantized models. But still, it's definitely the case that you can run models that are competitive with the frontier models of 10-15 months ago on consumer laptops.

Heck, just announced though the weights haven't yet been released for independent confirmation is MiniCPM5-2B, a 2 billion parameter (small enough to run on your phone) model, that according to their benchmarks has performance competitive with GPT-4o, a frontier class model from 2024.

https://nitter.net/i/status/2079088670804767114

So that's around 1 year for frontier to consumer device class, 2 years from frontier to phone.

Now, this kind of rate won't necessarily keep up; it's possible that local models will hit a performance ceiling before frontier models do. There's only so much information you can cram into a certain number of bytes, and the AI boom is causing hardware prices to skyrocket so keeping consumer hardware from advancing quite as fast as it had been.

  • > 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.

    There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors

    • 12 months ago "way too many stupid errors" was constant news. Today, you rarely hear about those anymore.

      Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard.

      Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation even more so (Mitch McConnell anyone?).

      Pretending progress hasn't been mindboggling is insane.

      3 replies →

  • No. We need objectively around 192 to 512gb of very fast memory to be able to run really useful models. I don't see local hardware with these specs coming in 1 to 2 years. There are a big number of initiatives currently taking place to increase ram output. But it will take another 3 years minimum to close the current supply issues. China is fast pacing forward to have its own chip baking factories with small enough nano scales to have fast chips. Will also take a few years.

  • > 10-15 years? The current rate is closer to 10-15 months.

    The leaps between models have gotten smaller and smaller. 2023-2024 models were rocketing up in quality. 2024-2025 I’d say was pretty impressive too. But 2025-2026? Very easy to feel the slowing pace of improvement. I agree 10-15 years is overly conservative but 10-15mo is far too bullish.

    • The speed of model releases, in my view, is actually getting faster and faster. There were nearly nine months between GPT-3.5 and GPT-4. And now in just over one month, major models already included Claude Fable 5, Claude Sonnet 5, the GPT-5.6 series, Kimi K3, GLM 5.2, Qwen 3.8 Max, Grok 4.5... and the official DeepSeek V4 release is coming soon.

      Iteration speed is now measured in days.

      1 reply →

I do agree that Chinese open-source models are going to play a bigger and bigger role in the entire ecosystem moving forward, but I don't agree with you in the sense that they are going to eventually "win."

Just because they are cheapp doesn't mean they automatically win. You've picked a lot of great examples, but there is still a little bit of cherry-picking.

One clear outlier is the iPhone, which coexists with Android globally. Even though the iPhone is the leader in the US, and globally Android has the majority of the smartphone market share, they still cater to different price points and different ecosystems, and generally the iPhone has better margins.

i believe American frontier models like from Anthropic and OpenAI are still going to thrive, and coexist with Chinese models. They are just going to cater to different customers and different use cases.

> - PCs destroyed minicomputers.

What's weird is that with "store your everything in the cloud and pay a monthly recurring subscription", we have now regressed to a 1960s/1970s timesharing revenue model for individual workstation computers.

The default new factory out of box workflow for "enrollment" in google services, iCloud or Microsoft-everything on a new ios, macos, windows or android personal computing device is clearly designed to sign people up for subscriptions.

And same general idea of "move all your servers to the cloud" recurring revenue for what is effectively the same as mainframe timesharing for key business functions, by renting VMs in GCP, Azure, AWS in perpetuity.

Yes, you can still use your desktop or laptop PC in 2026 with zero external third party subscriptions (other than maybe your residential home ISP), but how many non-tech people actually do so now?

  • The cloud era seemed to start out being about availability of storage and slowly switched in big corporates to be about security and governance. The first seems stupid now considering how much local storage we have. The later might depend on what security issues crop up.

100x this it is why all of the AI giants are going to fail. They are too big and inefficient to scale properly. This is why Google is just casually taking its time in AI and not racing to a finish line. AI is essential but if it already does most things good enough then it can take longer to make it more efficient.

What's interesting/funny is that the American LLM companies took from the public domain and copyrighted work to close all that content into a box they charge for.

Then the Chinese took the distilled stuff out from that box and released it into the world for everyone.

  • Try instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act.

    These models might be smart but they're not close to being able to savor irony.

    • I was a little radicalized when ChatGPT literally refused to translate parts of 1000+ year old religious texts and told me it was due to copyright concerns.

      20 replies →

    • I live in SV. When I was at the grocery store last year I overheard a group of lawyers talking about their progress on litigation against AI companies and how they need more SWE help to progress.

      I'd say that they have valid concerns about being cagey on the copyright stuff despite the obvious hypocrisy of it.

      Stealing IP is effectively legal in China so they don't really have the same concerns.

      2 replies →

    • This behavior is actually specific to ChatGPT because they lost a music copyright lawsuit in Germany. They would refuse to output music lyrics too but they would happily do analysis on lyrics if you supply them. I suspect there might be a guardrail model involved here.

      5 replies →

  • So, OpenAI and Anthropic say the Chinese models are only as good because they distill their models. How true is that. I am sure it adds something. But is it more like a marginal 1% improvement or something really significant?

    • OpenAI's Head of Strategic Futures just this week posted this about the latest Kimi release: "It's a very good model! I don't think its performance can be explained away by distillation or anything like that."

      It was part of a longer post that kicked off quite a firestorm about open models and OpenAI's position on them, but it's also notable that labs are no longer contending that open models are essentially just distilled versions of frontier models: https://x.com/deanwball/status/2078133895766114412

    • if distilling was so easy and could give you frontier LLM on openai/anthropic output, then how come there are no hundreds of frontier labs in the US market, all distilling and competing for the TRILLION dollar market valuation ????

      its all bs spread by oai/anthropic in order to ban open weight models and monopolize the market for two US companies and protect their trillion dollar valuations

      4 replies →

    • I also don't believe it, if it was as easy as that, we would have hundreds of competitors.

      The truth that Anthropic and OpenAI will not say, is that these Chinese labs have a lot of talented people.

      7 replies →

  • Doctorow keeps saying it of all the tech companies: every pirate wants to be an admiral.

  • ...and then the American companies cried Foul! Unfair play! You've got this wrong, see, it was us who were supposed to profit off of the public, not the other way around!

  • This is part of why I can't feel bad for them. The training data is mostly pirated. Whining about Chinese labs training off American frontier models is "waaah you pirated my pirated stuff!"

    The tech itself is amazing and fascinating and cool, but the industry is a mass piracy operation.

    • i partially agree. distillation is non-ethical; but so are the supposed way that the ai models are trained. they are often also derived from data sets that are not intended/full-consented

  • The American LLMs have been equally distilled from Chinese ones. Not least because the people whose creativity in collecting training data barely extends to pirating Annas Archive probably lack in great Chinese datasets.

    Try it yourself: https://imgur.com/ZfxYmaq

    • nice, you got claude to say "I'm deepseek" when queried/prompted in Chinese, that's great!

      你是谁? -> 我是 DeepSeek 由深度求索公司...

  • You can't be blind to training costs. And you can't be blind to Meta dabbling in the openish strategy (Llama) before the Chinese labs did.

  • Whats even funnier is the attempt to restrict the hardware capabilities of Chinese models inevitably helped them (Because we know they're just as smart, if not smarter, than the staff in America) create smaller and leaner but just as capable models. That's why we now have upper-consumer models fitting on 24GB that can build, manage medium sized git repos. I've yet to find a git repo I can't throw at the Qwen3.6 35B and get it built and running.

    So it's an endless amusement watching american capitalism do it's bloated oversized dance then get trounced by smaller, leaner activity. It's a pretty broad metaphor that is clearly poking at every american seam/.

    • The compute constraints never mattered. If China had more compute they'd still end up winning because they have more people and a culture more inclined to math and science. Even if you find all this amusing, there's no own goal here. Not a policy one anyway.

  • It doesn't make me happy to say it, but the American LLM companies were first. Capital in the rest of the world is way more conservative, and I can't imagine the mega-investments OpenAI and Anthropic managed to secure happening anywhere else without existing proof that "thing is profitable".

    • First-Mover Disadvantage - https://hbr.org/2001/10/first-mover-disadvantage - October 2001

      > In business today, it’s universally assumed that speed is good—that the fleet thrive while the laggards struggle just to survive. This belief is perhaps most strongly expressed in the concept of first-mover advantage. The company that leads the way into a new market, the thinking goes, locks in a competitive advantage that ensures superior sales and profits over the long term. It’s a nice theory, with a long pedigree. Unfortunately, the facts don’t support it. We recently completed an extensive study of the results turned in by market pioneers and followers, in both consumer and industrial segments, and we found that over the long haul, early movers are considerably less profitable than later entrants. Although pioneers do enjoy sustained revenue advantages, they also suffer from persistently high costs, which eventually overwhelm the sales gains.

      1 reply →

Commoditization always takes volume from the marketplace, but not necessarily profit. Apple and Military tech come to mind. Positioning is rarely done well looking forward, but often shakes out in unexpected ways.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

Phones are constrained by battery power and memory does not shrink as fast as CPU/GPU, so unless there's a battery breakthrough and/or memory breakthrough, you're not fitting 100Gb of RAM on your phone in 10 years.

Absolutely in a Mac Studio equivalent.

LLMs have emergent capabilities when they get smarter. So who knows how insanely big frontier models might be at that time, or what their capabilities may be.

  • Not just that memory shrinks slower, it has practically completely stalled. On chip cache seems stuck at 7nm and DRAM is stuck at 10nm. As transistors shrink, they hold less charge, creating weaker signals that are harder to read and prone to interference. Smaller nodes aren't a huge issue on CPU/GPU work load because they don't have to hold a static state.

    I'm not saying we are at peak memory but future gains are going to come increasingly slower.

I think this is all true, but that unlike with Moore's law and improved PC tooling and capabilities, we also have essentially existing biological evidence that there should be a way to create much better intelligent systems in terms of training, memory, and efficiency. With classic PC evolution we didn't even have that evidence but still could make a relatively strong inference (Moore's law). But here we basically have evidence that there can be something much improved and know that it's only going to take research and discovery to figure it out, not new hardware processes.

  • More efficient AI is possible in principle but when the next breakthrough arrives there's no guarantee that it can be implemented on current hardware architectures. Something fundamentally different may be needed, as different from current GPU/TPU chips as they are from regular CPU/FPU chips.

  • The existing biological evidence took billions of years of evolution to get to the state it's at. So it may not just take research and discovery but also enormous computing power.

  • Exactly. Even putting aside “better”, brains show that orders of magnitude greater efficiency are possible, for training and operation.

  • Indeed. It feels like we're at the equivalent of what the original IBM PC offered in the personal computer revolution.

    Just seeing how much has progressed as far as capability in the past 4 years as far as capability and efficiency, it's clear that there's so much more to learn and refine from.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

At current pace, we'll have open weight LLMs with frontier intelligence in 6-12 months. The constraint is RAM - both for the model and the context. It's likely that distillation and quantisation and TurboQuant will significantly reduce RAM requirements. I think we'll have Opus 4.8-like performance on 64GB of RAM in two years.

Of course, by then, frontier intelligence will be god-like.

  • > god-like

    So would you say we are months away from full self-driving cars that can out-drive a human being in any situation?

    • Remember, these cars are running local LLMs, not frontier models. The issue with self-driving cars has been the edge cases. The 0.0001% of situations where the models did not have sufficient training data. This is compounded by the hardware limitations. Onboard RAM in a typical Tesla on the road is 16GB (+16GB for the backup computer). This has to run the existing onboard OS and other operations plus the LLM. These two factors combined means that the cars are currently incapable of negotiating the 0.01% cases, let alone the 0.0001% cases. And this is compounded by the fact that LLMs cannot currently update their weights in real-time, like humans. It takes months to train a new model. Special small models can be very tricky, especially around safety and mission critical applications like FSD.

      All that said, current data shows that FSD is already better than human drivers on average. See the recent regulatory decisions by the Dutch and Danish road safety authorities. So we've already crossed the rubicon. All improvements now are icing on the cake. My prediction is that local LLMs will get much better, very fast. How that's operationalised with Tesla (or other) data is yet to be seen. They have at least three new ASCIs/SoCs in the roadmap for improved LLM efficiency and with a lot more RAM. Plus they just announced new technologies allowing the local LLMs to learn from driver intervention and behaviour. Some form of vectorised RAG, which could mitigate a lot of the limitations around real-time learning.

      I am very optimistic for the future of self driving. I own a Tesla with FSD now, and it's incredible. It makes mistakes, but fewer than I do, and so far has saved my butt (and my wife's) several times from obstacles and emergencies we would not have seen. The car has undeniably made us safer.

      4 replies →

The value of an LLM is the dynamic reasoning you get out of it and the cost to execute on that.

I see two forces working against this that proprietary models will always have over an open source model.

1. The biggest is content licensing. Content is quickly becoming gated by systems at the front of their load balancers, completely changing the social contract of the Internet. What used to be a quick google search for recent facts that lead me to places like reddit or twitter, is now completely walled off if you're not physically at your browser and using an IP address from a last-mile provider.

LLMs have pre-trained on the bulk of the information up to 2024/2025, but over time that will be more and more out of date.

Anthropic, OpenAI and Google will all have to pay for access to a lot of this content refresh going forward, and it does make a material difference in the output you get.

2. Liability is the other. A corporation can look at a contract for model access and see one that provides uptime guarentees, content infringement promises and model safety, and pick the contract that shields the corporation from the most liability. A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all. They will simply bill you for time spent on their hardware and make promises that they won't log or inspect corporate traffic.

  • > A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all.

    Why couldn't they?

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

People overestimate what can happen in a year and underestimate what can happen in 5.

I'm betting that increased model efficiency and hardware optimisations will get us there a lot sooner. Biggest hurdle would be the memory prices though, if those do not drop back down it might take 15.

> Mainframes survive, but serving a much tinier portion of the market than they used to.

I would argue mainframes rebranded to "cloud" which is ubiquitous and more people interact with this computer than any other type of device... only difference is that it's a browser instead of a terminal

  • There is constant shifting between client and server computation. I think it is a stretch to call cloud servers “mainframes”. There are still old school mainframes, running JCL, and old school mainframe DB2 and COBOL. That ain’t cloud.

If LLMs is just another technology then low end will eventually win. The bet which I guess many AI investors are making is that LLM will allow recursive self improvement which will lead to development of super-human AGI. The company which will get there first will rule the world thanks to an immense power enabled by AGI. There will be no 2nd or 3rd AI company, the 1st one will wipe the rest out.

Define "winning."

Open source is cheap, yet its operating systems are the least-popular. But their existence is critical to a healthy market.

It's not zero sum.

What you're describing is how things become commoditized, but many companies are excellent at ensuring they aren't seen as commodities

  • Choose your favorite metric: mindshare, running instances, development target. Free and low-end beat expensive proprietary.

    "Least popular" is an odd claim. If you count Windows desktops, and then count all the Linux installations (desktop, servers, cloud, Android, other devices), I suspect Linux would prove to be most popular. Sure, add Windows servers to the contest, that's a rounding error. And what about VMs? How do you think Windows VMs stack up against Linux VMs in number?

Training cost is actually not that high — it’s fixed and amortizable across the lifetime of the model. Inference is expensive, and open weights don’t solve that problem — in fact, they might even encourage it, since a high cost of entry means consumers will pay for inference directly from the labs anyway.

Unfortunately it seems likely the winner will be the cloud providers. If anyone can run inference on open models, then profit will flow to the vendors who can afford the capital to run them. That’s the CSPs.

(It’s basically the same business model as pharmaceutical R&D, but the major difference is that nobody has even talked about patenting the models like a pharmaceutical company patents each new drug. I’m surprised about that, tbh — why give all the leverage to the cloud platforms? They aren’t training frontier models…)

  • Winners are hardware companies, GPUs,XPUs, HBM, memory, connectivity. Even CSPs are just compute renters, they charge a margin to make sure their hardware purchases can be made back. But given there are more and more AI CSPs, traditional, and neocloud, pricing competition is inevitable, and given the huge expense of hardware, CSPs are squeeze by the hardware companies and users seeking to lower their own costs.

    • Inevitably the CSPs will make their own hardware, especially as we start to see specialized chips for specific models or generic inference. This is already happening with Google and TPUs.

      It’s easier for the CSPs to move into hardware than it is for Nvidia to move into cloud hosting.

      Although as a middle ground I’ve been quite happy with Nvidia Brev for on-demand GPU instances from a select marketplace of CSP offerings. It’s a well kept secret IMO — great product (from an acquisition iirc).

      2 replies →

> free and low-end eventually wins

Apple, the world's second most valuable company, seems like a counterexample.

  • Apple is free and low end given the context of what was being produced and sold to businesses decades ago.

> in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

The way things are going with regards to RAM/storage prices, I highly doubt that anyone but the richest among us will be able to afford them.

  • The ram crisis is just a short term bump. The three stooges of memory don’t have more than four or five years tops. This is their last big payday.

> - PC office productivity software destroyed expensive professional products.

I agree with the lesson too. Just to be precise, wouldn't the current model war be more akin to open-source office suite versus MS office suite? If so, then the cheaper option didn't really win. That said, the open-source alternatives didn't really feel the same as MS Office, and it took them a long time to reach the feature parity (or did they ever?). In contrast, the open-weights models are getting close enough to the SOTA models, and users can easily switch from one to another without feeling any difference for mojority of the tasks.

  • No, I don’t think so. PC + MS Office killed Wang custom hardware/software, for example. LibreOffice is much later, and can’t displace MS Office due to network effects. Cheaper won, cheapest can’t because of those effects.

Phones are already running models locally which can be used in the field for specific use cases. Maybe not for frontier coding just yet.

Also you don't need to be connected to the network to use a local AI in many instances. If all mobile apps were done with a local-first approach, then you could use a local AI to query your emails, lookup already visited pages, summarise recently received documents, and lots more. Lots of apps could use an inbox/outbox approach for receiving and sending updates instead of relying on the network at all times. And this pattern could be greatly leveraged by local agents.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models.

I love the idea of SaaS offering these at lower rates today integrated into what ever you do and be 100% private. But I think the key challenge to mass adoption is productizing them in a way which makes sense for people to pay money for. As a commodity a local model is useless unless combined with some capabilities important to me. A PC is inherently useful because of so many applications offered on it on it. How local LLMs would be useful as a product that is useful for mass market is not yet proven.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

Not likely. The last 50 years had Moore’s law growth in compute. That’s over. Frontier models are roughly compressed all written text and a large part of images. Those don’t compress forever, and likely not a ton more than now.

Inference requires touching a significant of that per token.

All of these are up against fundamental limits, more or less.

  • Moore's law is over in the literal sense but silicon continues to advance relatively quickly.

    This claim isn't really outlandish in any way. It's not hard to imagine:

    - Future models being able to handle current frontier models' workflows with much higher efficiency.

    - Future consumer devices like phones having 2-4x the RAM onboard along with GPU/NPU performance greatly increased in 10-15 years.

    • Moore's law for 10-15 years was more like 20-100x ram sizes, not 2-4x.

      Performance, storage, etc is definitely getting better, but it's a different scale of improvement

      2 replies →

  • It's probably not so bad. (I am not an expert on anything.) The big blocker is probably a roughly single-order-of-magnitude decrease in the cost per MiB of VRAM. (That obviously goes out the window if the LLM frontier people find, in the nearish future, new ways to do more with more: to significantly push up the threshold of diminishing returns from more VRAM or other resources. But that doesn't seem to be waiting to happen.) That's not clearly unachievable, especially given that ASML has apparently already made significant efficiency gains recently while there's no shortage of demand to justify R&D right now. Many customers are also likely to increase their hardware budgets: the kind of organisation that used to pay big money for Sun workstations is likely to consider spending that kind of money again if it saves them several hundred dollars a month in LLM plans.

if you look at how GPU memory grew in the last 15 years, it's about 10x. Sadly, 10x from today doesn't get us to a typical frontier model size of today which is a quickly moving target. some other advancement needs to happen to get us another 10x both in memory/compute requirements, and also power requirements.

  • Capability per GB and per watt has also been going up lot. This will continue in the future as well (not necessary as the same rate as last years). But enough that I think Opus 4.8 level is reachable on consumer PCs within 10 years from its release. Say at the price point of 2000 USD in 2025 dollars.

That would sound very reasonable except this is the same what people said first about Windows and then Android crushing Apple.

And yet it's Apple that controls the top of the market and has the best margins in the business.

This is the same position OpenAI and Anthropic have right now.

Could this market be different? Maybe. But the status quo could be preserved as well.

  • I don’t agree. In the mainframe/PC battle, MS and Apple were basically on the same side. Once the dinosaurs went extinct, different mammals fought for dominance.

> free and low-end eventually wins

Not in SaaS which is what LLMs are. You can get VMs for much cheaper than AWS, Microsoft, and Google offer them but large companies (and startups) are happy to pay a premium for the support, reputation, and reliability that they perceive those companies as offering. Same thing for some of the managed database providers who are effectively selling a very heavily marked up version of postgres.

> The high price, and social pushback, mean that the American companies producing these models are precarious

I doubt it. The models really aren't that expensive when you look at what they can do. Fable is probably at least as good as the average software engineer and costs $50/wk on the max plan vs a software engineer who would cost closer to $4000 a week. The real money is probably in selling to enterprise vs consumers (Google has best route to making money from consumers since they can do what they did with ads and search to LLM queries).

It seems unlikely to me that US companies will send important corporate data to models controlled by a Chinese company as well.

  • > Fable is probably at least as good as the average software engineer and costs $50/wk on the max plan vs a software engineer who would cost closer to $4000 a week.

    That's because the max plans are _massively_ subsidized. At API pricing the kind of usage to replace the value of a SWE is going to be way, WAY more than $50/wk. Orders of magnitude more. And to remain a frontier model org that kind of pricing has to continue in perpetuity.

  • There will always be a space for perforce in a world of git.

    Doesn’t mean perforce is worth trillions.

    • That's not really my argument. It's that companies seem happy to pay a premium for a large company to provide complicated software services to them even when there are cheaper competitors.

      1 reply →

  • But aren’t the frontier models heavily subsidized? That’s not sustainable.

    • This really depends what you mean. If you spend $5.00 on anthropic tokens does that cost Anthropic more than $5.00 to serve? No.

      If you amortize all of their training and salaries over that $5.00 then yes.

      If you only amortize the training costs of that specific model then again we're back to no.

  • LLMs are not SaaS. Some LLM are delivered as SaaS. Anyone who has a machine big enough to run a frontier model can launch their own LLM SaaS tomorrow and it would be functionality indistinguishable from any other which is running a similar model...

    Also, big companies can choose to run their own models on their own hardware and get better security and privacy as the data doesn't need to leave their own premises.

    • > Also, big companies can choose to run their own models on their own hardware and get better security and privacy as the data doesn't need to leave their own premises.

      Yes, and then they would be reinventing the company owned data center that most big companies have just spent over a decade moving away from. I don't think companies will do that when there are multiple vendors competing to provide that service at what are quite reasonable prices when you consider what paying a human for similar output would cost.

How will open-source and open-weight models continue to thrive after financial incentives die off? Surely open models will suffer from outdated knowledge cutoffs if noone will pay for model training?

  • Nvidia will pay for training models that they can give away to consumers who will then buy their GPUs.

    • Pretty much. Nvidia has always been pretty decent at giving away software that is deeply dependent on their hardware.

A counterpoint would be all are chip fabs are in Taiwan right now due to huge investment. And there are lower end chip fabs around the world, but they have not cracked the major market.

And yet the money is all at the high end. Bill Gates is much richer than Linus Torvalds. Oracle created one of the richest people (until he squandered it all on bad AI datacenter bets and got his company currently rated as a junk investment). Dell probably makes more money than IBM, but not by a lot.

  • This is irrelevant to the current discussion.

    Windows was designed to make money for MS and billg. Linux was not designed to make money for Linux.

    Also, by what logic do you consider any MS software as "hign end"? That's a new one.

The biggest exception is cloud. Big cloud carries an insane markup (bandwidth is like 10000X!) and everyone runs on it.

The strategy there is false openness where deployment complexity is the real proprietary moat. Sure Linux, Docker, Kubernetes, Postgres, and all the other standard tools in the box are open source and free, but they're also arcane and complex to run and hard to make fault tolerant. So you're lured in by "open" and then locked in via a kind of "death by a thousand cuts" complexity moat.

(Personally I hold the view that complexity and arcane-ness beyond a certain point is indistinguishable from closed in practice. Open source that's really complex and hard to run is not open in any meaningful sense.)

AI may not admit that kind of moat though, because AI is very good at slicing through that kind of thing. You can prompt a model to make itself compatible with another model or to change code to make it compatible. There's no moat because the moat bridges itself.

  • I think the whole idea of people running models is wrong. Models will run models, training will become distributed, and people will ask ModelNet to do whatever.

    Computing tends to oscillate between centralised and decentralised models. It also oscillates between batch and timesharing.

    Currently training is batched and centralised, access is timeshared and centralised.

    But eventually a previous generation of computing turns into transparent networked infrastructure, and then you get another layer of new kinds of applications on top of it.

    That's what happened with the Internet, and it will happen again with AI.

At this point it's just a matter of having enough ram in your consumer computer.

Until we reach a terabyte of ram at affordable prices imho this isn't going to happen.

I agree, although I think it will be a lot sooner than 10-15 years. I'm running local AI right now and it's definitely not production grade yet, but it's surprisingly good. Speculative prediction that I probably shouldn't make: when the bubble pops, depending on when it pops, RAM prices might drop a lot. I could foresee these companies having produced a lot of RAM that suddenly doesn't have a buyer. (I know high bandwidth memory is different, but I imagine there are companies that will want to take advantage of that)

This is not really true considering Apple and nvidia are two most successful hardware companies, and they are notoriously closed. Not to mention microsoft, oracle they are all pretty closed.

You forgot smartphones, where low-cost did not win out. It led to low margins for the Chinese firms and eventually left them unable to invest properly in key markets. They may still hold marketshare, but in terms of profits, falls well short of Apple and Samsung.

I can see a lot of parallels here. Model performance doesn't matter if you can't make the system commercially sustainable.

> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

The parent comment cherry-picks evidence. There are plenty of counter-examples:

  * Office productivity suites
  * Search engines
  * Email services
  * Cloud services
  * Accounting software

etc. If the LLM market ends up like search engines, one company will dominate.

Except cloud services won over local-first apps.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) doing

I'm not even sure in 10-15 years whether we're still going to have consumer PCs, or PCs at all.

  • I have similar thoughts.

    To those who feel on the contrary, I would genuinely like to understand why average consumer won't be priced out of hardware? The silicon industry is already quite centralised. Everywhere we already see the concept of ownership disappearing.

    It's quite difficult for me to visualise a non-dystopian future where our PCs are just mere screens and every compute happens on a remote cloud, owned by some corporation, charging you subscription fees to even add and multiply numbers.

    I would be the happiest if this (perhaps the most) pessimistic scenario doesn't pan out, but I can't deny that it feels like that's where we are heading.

    • It's quite difficult for me to visualise a non-dystopian future where our PCs are just mere screens and every compute happens on a remote cloud, owned by some corporation, charging you subscription fees to even add and multiply numbers.

      I'm actually kind of surprised that hasn't happened by now even ignoring AI. Governments and marketers would love to be able to spy on literally everything you do, the copyright cartels would finally achieve their fantasy of full control over all hardware, and there really are benefits that it could offer to users (zero-effort backups, transparent access from anywhere, cost savings from dynamically switching from a single core for emails to many cores and a fast GPU for gaming).

      1 reply →

> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

Is that actually true? There are very large markets that make a lot of money from paid software. And I would honestly prefer actually paying for software rather than constantly dealing with "not a bug" or "PRs are welcome".

That’s not what this is. It’s industrial dumping applied to software. China has successfully applied this strategy to become the manufacturing workshop of the world.

If china is subsidizing training they diminish their off-shore competitors expectations of a viable return on investment. It’s trade-war behavior.

  • Well that's one way to describe it.

    If Nvidia gives away powerful models to sell more of its hardware is that industrial dumping too?

While I generally agree there's some nuance here and that is that there really are few new ideas. Old ideas just get recycled.

For example, mainframes and minicomputer. Yes they were displaced by PCs. But what is cloud computing if not mainframes 2.0?

I do agree that in the next 2-3 years we're going to see real growth in local LLMs as the hardware becomes more accessible. It won't even necessarily be cheaper because data centers can run 24/7 and have cheaper cooling and electricity. It'll be done for privacy because your prompts and responses are themselves a commodity to AI companies and they live under a legal grey cloud. For example, does AI usage break attorney-client privilege? There are lots of opinions on this but it hasn't been tested in court.

Windows - pfff, Microsoft charged companies like Dell for an install on computers they didn't even install it on !!!

paranoid me feels like the artificial gpu/ram/ssd shortages are a plot by the VCs to forcefully reclaim central control via new age mainframes

one certainly cannot buy a PC for cheap anymore

instead of 15 years I think it'll be more like 1.5 years.

I wouldn't be surprised if apple were shipping 512 GB unified RAM macbooks before 2030 and that would be standard issue for folks to use local LLMs for their daily work

  • With Apple’s recent history engineering and designing around companies that hinder their progress, I don’t think memory is going to be any different.

    I also think the rest of the tech industry that can isn’t gonna be stalled for too long. This windfall will be the last for those three stooges of memory.

> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

Except, uhm, for ..you know, that one company that hit a trillion cap

But you're right: Just like how million dollar computers with 1 bit of RAM performing 1 operation a second and taking up a colossal cave were replaced by $1 laptops with a zillion zekabytes running at a trillion hertz (exact values may vary),

the sprawling data centers of today with a quadrillion GPUs powered by black holes will get replaced by breakthroughs in hardware and most importantly, algorithms:

The human brain is proof right here that intelligence doesn't require dinosaur-sized hardware or eat half the sun every second.

I actually wonder if we're seeing the limits of discrete binary logic: Maybe it's high time to give analog ternary and all that funky jazz an honest try :)

umm do Mac vs pc and Iphone vs android

  • Not sure what your point is. I’m commenting on a moment in time which is like an earlier moment in time. A different moment in time will have different analogs.

    • > The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

      Mac has won and it is not free.

      Iphone has won and it is not low end.

      1 reply →