Comment by tomr75

2 days ago

what token/s?

27 t/s. I suspect there will be significant speed ups in the coming weeks.

  • You should be getting way more than that on a 6000 pro even today. I'm getting 40tok/s on a pair of 3060s. You can ask a SOTA model to optimize your setup for you.

  • Getting 30 t/s on a Mac M5 Max laptop. You should be getting close to 200 if you can use the 5090 acceleration tools I see posted here. There are forks for the 4090 and the 3090. Maybe there's a fork for the 6000?

  • I'm seeing the exact same number on my Blackwell box. MTP put it at almost 70 which is pretty decent

  • Any idea why it’s so slow? the entire model should fit in the vram of one card.

    • The NVFP4 quant is completely broken, so I'm not shocked that other quants aren't fully there yet. Give it some time to cook.