Comment by BoorishBears
2 hours ago
Not surprising it's a hard cutoff: they almost certainly have two infrastructure configurations for the two max sequence lengths
Fewer nodes dedicated to prefill per instance, and fewer nodes in total since you don't need to support a higher KV cache.
Disaggregated inference also means they can tune the balance of compute dedicated to prefill seperately from decode
No comments yet
Contribute on Hacker News ↗