Comment by bel8
1 hour ago
true but if you're actually running k8s and similar workloads, chances are it might eat memory that LLM requires.
you'll also notice these articles rarely specify their context window in tokens, because it is small, usually 30k to 70k tokens and it gets slower as it fills up.
No comments yet
Contribute on Hacker News ↗