Comment by swiftcoder
8 hours ago
So the question becomes, how many other parts of the inference pipeline have left 1000x optimization opportunities lying on the table?
8 hours ago
So the question becomes, how many other parts of the inference pipeline have left 1000x optimization opportunities lying on the table?
The problem with the rest of inference is that changes are not trivially correct or incorrect, as they are with the tokenization layer.
Some changes certainly can be. If the model produces the exact same output for a fixed seed across a variety of inputs after a code change, I think it's reasonable to expect that the change is correct. There are also mathematical transformations that can be applied in some cases that are provably correct. (Not suggesting there's necessarily anything of this nature that will lead to 1,000x improvement though.)
Eh, linear algebra changes are still easy to measure correctness, it's just that you're competing with 50 years of research for most of them, less low hanging fruit.
the answer is many! This would take hours to write. Full teams and research on nearly every part. So many 'unlocks' coming.
I'm sure there's been a lot more effort put into the other, more consequential, portions of inference time.