Comment by toebee

1 hour ago

We got a rtx 4090 handling around 10 concurrent requests at 50 ms TTFA after some config changes / adjustment as it doesn’t have FP8. So this 50 ms TTFA thing is very much possible on consumer hardware.

I'm going to see how this shakes out on my machine with a few 3090s. I see you all are leveraging some custom cuda kernels, so it may not work out of the box on Ampere (30xx) architecture yeah?