Comment by crubier
18 hours ago
Came here to say that, my bet is that in 3-4 years you'll be able to run Fable-level of intelligence models on your laptop or maybe even on you phone
18 hours ago
Came here to say that, my bet is that in 3-4 years you'll be able to run Fable-level of intelligence models on your laptop or maybe even on you phone
But isn't there the raw intelligence of a smart model and then the practical intelligence fuelled by how many parameters it has? You probably will barely be able to fit a 70 billion parameter model on a phone in 3-4 years let alone a 2+ trillion parameter model... so it depends on what you call intelligence
I'm not willing to believe in phone-based frontier models anytime soon. Though, Gemma 4 12B is a beast that runs comfortably on the current top of the line phones (or would run fine if allowed to run, I think there's some kind of 6GB limit on iOS, and 12B is ~7GB). I'll believe in three years we'll be able to run ~30B models on the best phones. That's 16GB in a 4-bit quantization, and I believe ~30B models will be competitive with 120B models of today, based on the curve we've been on. Qwen 27B and Gemma 4 31B are competitive with much larger models of a couple years ago.