Comment by SV_BubbleTime
3 hours ago
There’s a lot about this that would make sense.
But, not really at the current technologies. Kimi and GLM are fucking awesome, but I don’t have 3TB of VRAM to run them, and I don’t expect to even when ram prices drop.
So now you’re back to the scaling issue before talking about power and compute distribution.
Do you think you'll realistically need 3 TB RAM to run a sufficiently good model 1 to 2 years from now? I certainly don't. Considering what can already be done with 128 GB of relatively slow unified memory, imagine if efficiency improvements continue apace, the memory becomes 128 GB of HBM, the flash device becomes capable of sequential throughput matching today's DDR5, and such a system was affordable as a routine purchase for the average person.