← Back to context

Comment by bel8

15 hours ago

Local inference solves that.

We used to pay for phone calls not long ago. Now a video call is free.

This doesn’t make sense. You mentioned OS and drivers, why send the commands to your hardware through the black box and extra hoops instead of directly doing so?

Do you believe the average person can afford the 1TB+ VRAM required to run a competent model?

  • The expectation is of course that that will not be required. Competent models will become smaller and the compute to run them will become cheaper, once the market stabilizes. That will take some time of course.

  • Eventually yes. You can buy greeting cards today for $5 that have more compute power than the Eniac.

  • Not today. But just like the average person can now afford a smartphone which is more powerful than a workstation not long ago, we will have enough computing power for good enough local inference, at an accessible price.

    • We're not in the 2010s anymore, consumer hardware is actively regressing. Smartphones are actually a great example to bring up, because I believe 2026 was the first year in which smartphone hardware did not advance at all (and arguably declined) relative to price, due to AI-related shortages.

      Nobody can predict the future, of course, but I personally believe this pattern will hold. A decade from now, the idea of a consumer being able to purchase a personal device with more than 8GB of RAM will be a thing of the past, all significant compute will be done in the cloud using the massive amounts of hardware being hoarded by corporations as we speak.

      2 replies →