← Back to context

Comment by SideQuark

16 hours ago

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

Not likely. The last 50 years had Moore’s law growth in compute. That’s over. Frontier models are roughly compressed all written text and a large part of images. Those don’t compress forever, and likely not a ton more than now.

Inference requires touching a significant of that per token.

All of these are up against fundamental limits, more or less.

Moore's law is over in the literal sense but silicon continues to advance relatively quickly.

This claim isn't really outlandish in any way. It's not hard to imagine:

- Future models being able to handle current frontier models' workflows with much higher efficiency.

- Future consumer devices like phones having 2-4x the RAM onboard along with GPU/NPU performance greatly increased in 10-15 years.

  • Moore's law for 10-15 years was more like 20-100x ram sizes, not 2-4x.

    Performance, storage, etc is definitely getting better, but it's a different scale of improvement

    • 10 to 15 years from now the scale of LLM efficiency improvements is, quite literally, unpredictable.

      It could be that the company valuations crash tomorrow, and (almost) only performance gains achievable on hobbyist-level hardware come to fruition from there on out.

      Or it could be that in the future, we have a custom "model FPGA" à la Taalas [0] in every home, and that it turns out we can still massively boost inference efficiency due to novel discoveries like TurboQuant [1] or a somehow-improved quantization method [2] again and again ten times over.

      Point is, Moore's law in this context shouldn't be applied to just hardware spec sheets alone, but more the total number of "parameters potentially improving", IMO.

      [0] https://chatjimmy.ai

      [1] https://research.google/blog/turboquant-redefining-ai-effici...

      [2] https://prismml.com/news/bonsai-27b

It's probably not so bad. (I am not an expert on anything.) The big blocker is probably a roughly single-order-of-magnitude decrease in the cost per MiB of VRAM. (That obviously goes out the window if the LLM frontier people find, in the nearish future, new ways to do more with more: to significantly push up the threshold of diminishing returns from more VRAM or other resources. But that doesn't seem to be waiting to happen.) That's not clearly unachievable, especially given that ASML has apparently already made significant efficiency gains recently while there's no shortage of demand to justify R&D right now. Many customers are also likely to increase their hardware budgets: the kind of organisation that used to pay big money for Sun workstations is likely to consider spending that kind of money again if it saves them several hundred dollars a month in LLM plans.