Comment by toebee
1 hour ago
We do graph capture etc at startup (same as vLLM) but this model variant doesn’t require prefix caching - the prefix is just 10 tokens.
1 hour ago
We do graph capture etc at startup (same as vLLM) but this model variant doesn’t require prefix caching - the prefix is just 10 tokens.
No comments yet
Contribute on Hacker News ↗