Comment by scottcha
8 hours ago
I run an AI platform and we need to tokenize fast and early to make a lot of decisions on the subsequent steps (things like routing, rate limiting and such). Its really important to do this efficiently even though its not a large % of total end to end time for the request.
To concur it's "latency critical", not "performance critical", people often confuse those two - optimize it all, but especially the chained critical path latency!
Latency isn't performance? Maybe you mean "not throughput-critical"?
but according to Little’s law, if you improve latency, you also improve throughput, right?
If I have the same number of CPU cores and they all can do their work in half the time they can double the number of requests now
1 reply →
Same here, as we're sitting in the middle between requests and what budget constraints are allowed given a particular token allowance there can be 10 ~ 100 milliseconds improvement in the UX (TTFT) given such massive tokenization speed up.
‘I run an AI platform’ I have so many genuine questions I don’t even know where to start.
Pick one. I would like to hear it.