Comment by ford
17 hours ago
Also kimi 2.6 at 1000tps (as of may), though when we reached out they had a >12 month waitlist and minimum 7-8 figure annual token spend.
[0] https://www.cerebras.ai/blog/cerebras-kimi-k2-Enterprise
17 hours ago
Also kimi 2.6 at 1000tps (as of may), though when we reached out they had a >12 month waitlist and minimum 7-8 figure annual token spend.
[0] https://www.cerebras.ai/blog/cerebras-kimi-k2-Enterprise
7-8 figures annual spend will buy a hell of a lot of capable local inference hardware you can own, though it won't be at the absurd token/s rate, you'll be able to run almost anything on it... And it'll still have a good residual resale value after 4 years the way things are going now.
Try and spend 10 million dollars on GPUs and see what you get quoted for lead time
Feel like you could spend 6 figures building out a team and the rest renting compute for a whole year, and get the team to create a local inference solution with that kind of budget…
yeah. K2.6 can run on insane speeds. So sad that they don't have K3 yet.
But it can apparently also run 5.6 Sol
Yeah but at “call to discuss pricing” rates