Comment by cesarvarela
5 hours ago
It's funny that most of the AI industry is built around the assumption (which is most likely true) that it is not possible to run SOTA models on current consumer hardware.
Imagine if someone managed to run an Astra- or Fable-level model on a 5090 at reasonable speeds.
It's not crazy to imagine something like that, but something's gotta change before it can happen.
Either the 5090 part - new hardware that's tuned for AI specifically. But we won't see that until the datacenter buildout collapses or finishes, since they are buying up all of TSMCs capacity.
Or perhaps it comes from the model. 1-2 years ago it would be inconceivable to use a 27b model for coding and expect any kind of usable results. Today, I have a model that feels like it crosses the threshold from a toy to a tool, and i can run it on dated pro-sumer hardware. I don't think we'll ever see SOTA on consumer hardware, but as the small models cross more and more thresholds the gap will matter less and less.