Comment by narrationbox
5 hours ago
Haven't read the full report yet, just a quick question. Are your numbers for cold start without pre fill or is it after warmed cache?
5 hours ago
Haven't read the full report yet, just a quick question. Are your numbers for cold start without pre fill or is it after warmed cache?
We do graph capture etc at startup (same as vLLM) but this model variant doesn’t require prefix caching - the prefix is just 10 tokens.