You should be getting way more than that on a 6000 pro even today. I'm getting 40tok/s on a pair of 3060s. You can ask a SOTA model to optimize your setup for you.
Getting 30 t/s on a Mac M5 Max laptop. You should be getting close to 200 if you can use the 5090 acceleration tools I see posted here. There are forks for the 4090 and the 3090. Maybe there's a fork for the 6000?
27 t/s. I suspect there will be significant speed ups in the coming weeks.
You should be getting way more than that on a 6000 pro even today. I'm getting 40tok/s on a pair of 3060s. You can ask a SOTA model to optimize your setup for you.
Getting 30 t/s on a Mac M5 Max laptop. You should be getting close to 200 if you can use the 5090 acceleration tools I see posted here. There are forks for the 4090 and the 3090. Maybe there's a fork for the 6000?
I'm seeing the exact same number on my Blackwell box. MTP put it at almost 70 which is pretty decent
Any idea why it’s so slow? the entire model should fit in the vram of one card.
The NVFP4 quant is completely broken, so I'm not shocked that other quants aren't fully there yet. Give it some time to cook.