Comment by bjackman
9 hours ago
I agree but worth noting that it's never gonna be very practical to run LLMs like this at home. Unless we have some sort of design breakthrough, the only "sensible" way to run them is at high batch levels on shared HW.
Like, yeah if I could spend a few grand on such a GPU I probably would coz I'm a rich nerd, but I'd acknowledge it as an extremely inefficient luxury, kinda like a sports car.
So I think you could say the real misfortune is that we don't really have the technology (be it computer tech or political/social tech) to do that shared-HW thing in way we can truly trust.
We could make LLM inference 100x cheaper to run at home efficiently, but that solution might need to be updated every 1-2 years, whereas current GPU are useful for various others tasks and last longer
"Never" is a long time. Just think about how much ram we had 10 or 20 years ago. 1.5TB isn't a lot really.
It doesn't matter if you have the RAM, running a 1.5TB model for a single context stream is fundamentally inefficient.
The typical ram has surprisingly not increased very much in 10 years.
> April 2016, 8 GB was standard across the 13-inch MacBook Air range
... Now it's 16.
Rich nerds will have quite a bit more. But I suspect the standard of model rich nerds want to use will have gone up somewhat too.
‘Never’ is a big word in the computing world. 10 years from now a model this size will probably run on a high-end phone.
Of course, by then we’ll want to run something commensurately larger.