Comment by bel8

1 hour ago

true but if you're actually running k8s and similar workloads, chances are it might eat memory that LLM requires.

you'll also notice these articles rarely specify their context window in tokens, because it is small, usually 30k to 70k tokens and it gets slower as it fills up.