← Back to context

Comment by gerdesj

11 hours ago

I (we) run Qwen3.8-27B-FP8 on a DGX Spark box - that's roughly £4000 of hardware.

I did benchmark it in various ways and it runs quite well but it is a quantised jobbie and 1.5k t/s is also rather faster than anything I can possibly hope to achieve.

To run that model at those sorts of speeds is going to need some serious investment and you are going to have to pay for it.

The problem is most providers hit tok/sec limits really fast. 1m/min is the default and the only place I can get 10m+ is from first party providers without a lot of upfront cash.

How fast is it?

  • Parent already responded but just for reference an RTX 5090 with Ninfer hits 160 tokens/second with qwen 3.8 27B which is very usable.

  • With MTP and FP4 I max out at 30ish t/s on mine. Without MTP or in regimes where the drafter performs poorly it’s about 10 t/s. FP8 is about half that