Comment by ipsi
6 hours ago
The future for whom? The general public? Not a chance, no way, not unless it's able to run on a phone (anywhere from 20-40% of internet users, world-wide, are phone-only).
For companies? I think that's a lot more plausible, as that's mostly just a question of money - is it cheaper to run and administrate our own models, or outsource that?
For technically inclined users? I think that's unlikely unless they're able to operate on relatively cheap hardware while still being just as good as the hosted models. And by that I don't mean "a mac studio," that's far more money than I think is reasonable. A single RTX 5080, maybe, once memory prices start to drop.
Compare the games your average high-end smartphone can run to the AAA titles of the 2010's. It's not a matter of "unless it is able to" but "when it is able to".
That's going from 150w 720p gaming to ~15w 720p gaming in ~10 years. Let's say an inference cluster draws 1500w to deliver a small-ish 500b model at reasonable speeds/quantization.
Extrapolating from your gaming example, it will take smartphones only... *checks clipboard* ...100 years to achieve datacenter-level performance at the pace of 2010's improvements.
Uh, my clipboard says 20 years