Comment by gerdesj
10 hours ago
128k context is not a limit of the model, that's a limit of implementation:
"Context Length: 262,144 natively and extensible up to 1,000,000 tokens."
10 hours ago
128k context is not a limit of the model, that's a limit of implementation:
"Context Length: 262,144 natively and extensible up to 1,000,000 tokens."
We're talking about the Cerebras implementation, which is limited to 128K.
It's in the link.
TPM means Tokens per Minute.
GP is referring to GGP’s last paragraph. 150k t/m, yes, and 128k context.