Comment by efficax

5 hours ago

all the effort in this space is going into multiuser, datacenter workflows for high throughput inference. running in resource constrained environments is not where the money is. but it will be, especially if we look forward to a world where having 512gb unified ram is normal for "workstation" machines. The semiconductor space is slow enough to respond that it's likely it will take a few years before production capacity has ramped up enough to get us past the current supply crunch but it seems inevitable to me that we'll be able to run huge models like kimi 3 locally in the next few years (maybe 2029/2030 for it to really be affordable)