← Back to context

Comment by zackangelo

14 hours ago

We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).

https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.

I tried in your playground and got 14.2 tok/s?

I don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?

any way to see the tok/s for all the models listed on your homepage? curious which has the best speed/quality tradeoff for me