Comment by 0xcde4c3db
3 hours ago
I don't think "plateaued" is the right word, but I do feel like there's been something like a logistic curve compression in the difference between smaller and larger models as the field evolves. For inference at least, the scale of practical difference between a single high-VRAM GPU or SFF UMA box, a whole rack, and a whole data center seems to be falling far short of what we might have imagined just a few years ago. The conversations I've heard have largely turned away from breathless anticipation of the next frontier model and toward attempts at hard-nosed evaluation of which tokens are worth the cost.
I think it’s more that pushing frontier is extremely costly and there is no free lunches in same way as 2024.
Maybe, though that's another kind of progress in itself. Very impressive progress!