Comment by jasongill
14 hours ago
It would be great if they made their inference capacity for this model available via OpenRouter; the fastest provider on OpenRouter right now is at ~80tps https://openrouter.ai/qwen/qwen3.8-27b#providers
They do appear to host other models on OpenRouter so maybe Qwen3.8 will be there soon: https://openrouter.ai/provider/cerebras
We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).
https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.
I tried in your playground and got 14.2 tok/s?
apologies we just got a sudden burst of new users and traffic, it's scaling up now.
just added 8 more H200s to the cluster, if you (or anyone else) runs into issues please feel free to drop me a message: zack at mixlayer.com
4 replies →
I don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?
any way to see the tok/s for all the models listed on your homepage? curious which has the best speed/quality tradeoff for me
This feels great