Comment by zackangelo
12 hours ago
We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).
https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.
12 hours ago
We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).
https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.
I tried in your playground and got 14.2 tok/s?
apologies we just got a sudden burst of new users and traffic, it's scaling up now.
just added 8 more H200s to the cluster, if you (or anyone else) runs into issues please feel free to drop me a message: zack at mixlayer.com
Works much better now! Got 103.9 tok/s, not quite 200 - but still amazing! Thanks for sharing
1 reply →
FYI, I might be missing something but I think your billing system might not be working well - I'm not seeing any indication in the UI that my usage is being deducted from the $5 of free credits.
1 reply →
I don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?
any way to see the tok/s for all the models listed on your homepage? curious which has the best speed/quality tradeoff for me
This feels great