← Back to context

Comment by MisterKent

2 months ago

Apple is actually interesting. They are one of the few companies with a chip / PC play with real power AND basically no play I'm the hyperscalar market.

That means they're actually incentivized at least short term, to benefit PCs becoming strong enough to do local LLMs. Which makes this play make even more sense. Though, I've been saying for a while that the local AI inflectiom point is the death knell for these frontier labs.

> Though, I've been saying for a while that the local AI inflectiom point is the death knell for these frontier labs.

"Death knell" is a touch hyperbolic. Hardware that can only run quantized models that take up GBs in VRAM falls short of even an A100 (by almost an order of magnitude[0]), which in turn falls short of what an 8xH100 cluster can do (also by another order of magnitude[0]).

I'm an avid believer in local LLMs, but I cannot deceive myself - data center accelerators will win on power dissipation numbers alone[1], even when giving generous allowances for higher efficiency on Apple chips - and assuming the Apple-efficiency advantage persists on the same TSMC process node.

0. Based on my unscientific fine-tuning training experiments across local and rented GPUs. YMMV for inference.

1. Unless Apple surprises everyone and brings back the XServe with M7, if not, then laptop and desktop for factors simply can't dump heat fast enough to compete head-to-head, and will be designed for lower input wattage.

  • Doesn’t need to be a winner head to head. If it can do 90% of the tasks the big boys do, at 50% speed, for virtually no extra overhead cost save for the power consumed by a prompt - that’s gonna work for a lot of people. And that’s also basically where we’re at today. Qwen3.6 35b running quantized on 10 year old hardware solves basically all of my uses cases for agents except for coding.

    The frontier models are faster, and better at coding, but not so much that i’ll pay $200/month for them.

    • Consider this. One of the smallest Qwen models (4B parameters) powers my home automation voice assistant, and runs on CPU alone at >20 tok/s. It is enough for that use case, and could be made even better/faster with a modest GPU. It isn't as smart as some cloud-connected thingamajig, but I would never allow a literal Google or Amazon bug in my home. Huge SOTA models aren't relevant everywhere. Most people use LLMs for rather trivial tasks such as finding typos or drafting text.

      5 replies →

    • > If it can do 90% of the tasks the big boys do, at 50% speed

      I want to live in this world too, but these numbers, as of today, are very aspirational and far removed from reality.

      I'm no tokenmaxxer; I find my modest local setup useful, I also know the limitations, it's slow and it sucks (relatively) at high-level and/or long-context planning, compared to frontier models. Only a minority of my prompts are max-effort - its not all I do, but, it also means frontier labs aren't dying any time soon

      10 replies →

    • This is what makes sense for me as well. All I need a local model is for playing with simple graphics: no gradients, at most ten colours which I can push through VTracer to get an SVG. Draw Things does the job, usually in 120 seconds or less.

      Sometimes, I need a quick throwaway bit of python. That can take 30 minutes of my time.

  • We'll likely see a transformation in how frontier models are trained as a result of a push towards local inference. While it seems unlikely now, given current pricing for RAM, in 10-15 years it's not unthinkable to assume we could see individual machines with 10-12TB (and well beyond that) of RAM which are accessible to the GPU. Min/max system RAM increased a LOT from 2010-2025 and largely because it was cheap. Once the hyperscalers aren't generating revenue for the RAM manufacturers, I wouldn't be surprised to see a massive push towards consumers in order to maintain gross profit. Not to mention new players who enter the market because the margins are measurably absurd right now.

    At some point there will be diminishing returns towards the "just throw more RAM at it" approach the current frontier models are taking. Commoditization is just as inevitable as it ever was... and in doing so will enable actual leaps of what AI/ML is capable of. That's not to say there won't be a place for 99.999999% accurate vs 99.99999% but those cases will be limited and likely prime to disruption based on real innovation vs access to capital.

    • The 1080ti is out there for almost 10 years now. It has 11GB of VRAM. A 5090 has 32GB.

      SOCs with unified memory have shifted this a bit forward, but they're also expensive as shit.

      10TB ram in a consumer device is simply not happening in the next 10 years.

      3 replies →

    • I agree with the general direction but I'm a little skeptical of the "just add a few more TB of RAM and the frontier moves local" version of it

  • The established AI players have no financial interest to make LLM available locally. They aren't hardware companies and if running LLM requires paying them to host the models as well then they can naturally capture more of the value chain = more revenue.

    Apple is the only player here where it would play into their natural hardware incentive to get you to pay more for better hardware. It would make sense for them to find a way to run LLM locally (eg, newer architectures that others here have pointed out).

    Interesting times.

  • Is it hyperbolic though? One of the best things about the compute and memory shortage is that people are going to insane lengths to optimize things to run on lower memory / lower compute devices. If we keep this up for a while and then ramp up memory and local compute production, that AI inflection point may actually come.

    Of course, these are a lot of ifs.

  • The big question for local LLMs is whether there is a 100 tok/s model which requires less than 16 GB of memory and is competitive on most tasks with the cloud models.

    There is some signal that this is possible through both hardware innovation and training/data improvements.

    Cloud models have their own constraints - I can’t have opus4.8 spend 4 hours on a deep research question I had in the shower without spending money. I can’t do real time video game upscaling and graphics work in the cloud period.

    A laptop is about an order of magnitude cheaper than a cloud server thanks to economies of scale, uptime requirements, and other factors.

    • > The big question for local LLMs is whether there is a 100 tok/s model which requires less than 16 GB of memory and is competitive on most tasks with the cloud models.

      Benchmarks maybe? Real world, no.

      You just need the context otherwise. There's no way around it.

      1 reply →

    • if you do the electricity math you'll see that you pay more on local models while getting less (local is more heavily quantized) compared with OpenRouter.

      I'm not talking local Gemma/Qwen vs cloud Opus, but against OpenRouter same Gemma/Qwen

      there are reasons to run local - privacy, availability, but cost is not one of them

      8 replies →

  • The thing is, with the level of hard investment AI vendors have, even a small reduction of their addressable market is significant. They aren’t profitable, and inference is getting commoditized fast, so even if they eventually become profitable (not via financial engineering) they won’t be able to have good margin. The pressure of both open models AND local models is pretty bad imho

  • > Hardware that can only run quantized models that take up GBs in VRAM

    That's the today hardware.

    Now suppose Apple goes to any of Samsung/Micron/Hynix and says "we'll pay you the entire cost of building another DRAM fab and in exchange we want its entire output" and then releases M7 devices with enough memory and compute to run bigger models.

    > Unless Apple surprises everyone and brings back the XServe with M7, if not, then laptop and desktop for factors simply can't dump heat fast enough to compete head-to-head, and will be designed for lower input wattage.

    Laptops maybe. Desktops can dissipate more heat than the amount of electricity you can draw from a typical household wall outlet.

    • > Now suppose Apple goes to any of Samsung/Micron/Hynix and says "we'll pay you the entire cost of building another DRAM fab and in exchange we want its entire output"

      It's revealing that they aren't doing this: no one wants to fund that gamble on the state of AI demand 12-18 months out, but ate happy to capitalize on their current product lines/capacity.

      > Desktops can dissipate more heat than the amount of electricity you can draw from a typical household wall outlet

      100% agree, but the data center power and cooling infra are not limited by home wiring, and go way beyond what a wall outlet can safely provide (1,440W max on a typical 15A circuit at 120V). A single H100 maxes out at 700W

      1 reply →

  • I'm not paying for a super computer to do my taxes if a cheap pc can do it for free.

    So yeah, commercially it might be a death knell. Yes there's still a market for super computers, but would your rather own Apple or Cray?

    • > would your rather own Apple or Cray?

      I would consider an HPE tower server with a processor on the same league as an M6 or M7 under the Cray brand.

  • Indeed. Local models becoming available and halfway decent don't obviate the laws of scale. And because there's no ceiling to what scaling more will buy you in terms of capability, there's no reason not to scale more, there's no incentive for billionaires not to grab all the fab capacity they can.

    Enjoy paying $1000 or more for a little 4 GiB cloud terminal that connects you to all your online accounts where all your actual work gets done. This is the future.

    • >there's no ceiling to what scaling more will buy you in terms of capability

      This is highly doubtful.

      Rule of thumb: everything people think is exponential is actually an S curve.

      3 replies →

Indeed. If Apple makes it feasible to run models like GLM 5.2 at home, I will become their customer.

  • It's plausible but is the Apple Tax for a 1TB memory machine on top of current memory prices really worth it? I paid around $4000 for 4090m laptop with 16GB VRAM back in 2023, it's great but DoA for even quantized LLMs. I can run SLMs and fine tune it but that's it.

    We need one of those specialized inference chip startups to succeed and a PC manufacturer willing to bet on them against Nvidia for the local AI to find mass market appeal.

    • I recently bought a Mac mini M4 16 GB - mostly to run Immich. I assumed I needed a Linux box. After a lot of researched I was quite surprised that the mac was the cheapest option. So not always an Apple tax.

      13 replies →

    • > I paid around $4000 for 4090m laptop

      That's how much many developers currently spend on tokens - every day. Whatever "Apple Tax" applies to a device that can run a capable model offline will amortise itself in a blink.

      4 replies →

    • "Apple tax" is such a lazy and inaccurate accusation to level. Sure we've had expensive wheels on the cheese grater (ie Mac Pro) but we've also had:

      1. When Apple came out with the real Macbook Air in 2010/2011 (not the silly 2008 one), nobody could compete with it with those specs at that price and they couldn't for years. And every competitor usually sucked in some major way, most often the trackpad;

      2. The Mac Mini is an outstanding piece of hardware for $600. Or was;

      3. I've generally found that "Apple tax" complaints levelled against the iPhone to be nothing more than Android cope;

      4. The M-series silicon has been an absolute game-changer. I honestly thought the first-generation M1s would be not great but they came out swinging. And the price points for these Macbooks have all been great, much better than the last-gasp-of-Johnny-Ive touch bar butterfly keyboard series, which were objectively awful.

  • you're not a customer of any of their products at all already? not a single apple device in your household?

    • I didn't have a single Apple device in my house until a month ago when I bought a Neo. The last Apple devices I had before that were an iPod Nano and a PowerMac G5 many many years ago.

      Apple has pretty good competition in every segment with the exception of maybe the iPad, but I'm not a tablet user.

    • No, I've never owned an Apple device in my life, neither has anyone in my family to my knowledge.

    • None. And I have a PC, a personal laptop, a work laptop, my current and my previous Android phone.

    • What a bizarre bubble you live in to even be asking this question... I've never owned a single apple product, and never will.

      And in the rare occasions in which I have to use someone's MacBook, I'm completely lost - like some elderly person.

    • there a many people who don't own Apple. Why are you so surprised? I certainly don't and never will. What's it got that I can't get on a standard PC + Linux?

Tangential: About 8 years ago ex-Apple chip engineers left to design server-grade chips, this was Nuvia, and they got sued by Apple to the point that they had to get acquired by Qualcomm.

I worked at a hyperscaler when the M1 came out. A MacBook Air M1, running a Linux VM was faster and more energy efficient than anything we had in the data center.

  • I'm certainly willing to believe the M1 was the most energy efficient, but your data center didnt have anything faster than a laptop?

    • Nope, across all benchmarks, Linux running in a VM on the M1 had higher perf (4 cores) than any instance type I launched across Intel, AMD and Arm. Triad, coremark, openfoam, verilator, etc. This may have changed, but my hunch is this is still probably true with M5 vs all the cloud providers.

    • That's not especially surprising for low thread counts. e.g. if you look at the current Geekbench single core charts, the fastest device across Android/iOS/Mac/PC is the M5 MacBook Pro. Second-fastest is the M5 iPad Pro. The fastest PC CPU (9950X3D) is about 80% the performance of the M5. (And the 9950X3D is about on par with the A18 in the iPhone 16 from 2024.)

      (All the usual caveats about geekbench scores apply, but they're not nothing.)

They do stand in front of a great opportunity that would also benefit consumers, which seems rare in the llm era.

If people can get opus4.6/gpt5.5-like models locally, labs could raise their prices and sell token speed, better reasoning, mobile-focused improvements, you name it.

Not all consumers are power users and many will be happy to pay for flexibility.

>AND basically no play I'm the hyperscalar market.

I do wish they have Xserve back, or a Mac Pro that is Rack Based and support multiple node with M6 Ultra. The Hyperscaler market is so large along with AI their old business case of Xserve didn't make sense no long hold true.

I really wish people stopped saying things "I've been saying that"

why not just say "I think that"

do you see yourself as some kind of visionary about this particular topic? literally EVERYONE is saying that, it's the most obvious fact about AI

I'm not sure it's a death knell for frontier labs so much as a narrowing of what people need them for

  • When you've raised hundreds of billions in funding, every result except "to the moon" is a death knell.

    • LOL at the downvotes. I'm sure that the shareholders aka wealthy people at the top have a lot of patience for huge gambles that don't pay off. If OpenAI and Anthropic turn out to be the IT equivalent of mom and pop stores 5 years from now, their current shareholders will be ecstatic!