Comment by jaggederest
6 hours ago
To concur it's "latency critical", not "performance critical", people often confuse those two - optimize it all, but especially the chained critical path latency!
6 hours ago
To concur it's "latency critical", not "performance critical", people often confuse those two - optimize it all, but especially the chained critical path latency!
Latency isn't performance? Maybe you mean "not throughput-critical"?
but according to Little’s law, if you improve latency, you also improve throughput, right?
If I have the same number of CPU cores and they all can do their work in half the time they can double the number of requests now
I don't think that's accurate. If tokenization takes say 10ms and the rest of the inference steps take 50 ms then, improving tokenization will improve the time to first token but won't affect throughput much. After the first token, the inference steps effectively hide the tokenization time.