Comment by KerrAvon

3 hours ago

We're on the cusp of Kimi K3 becoming usable on sub-10k hardware.

https://github.com/gavamedia/deltafin

14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp!

  • It's need it to be an order of magnitude cheaper, but for some queries ("What to discuss at tomorrow's meeting") I can wait 12+ hours.

  • At that rate, it would take me a mere 15 days to generate the number of tokens I typically use in a day.

    "Yes, I'm using AI to speed up development. I'll submit that PR in two weeks time!"

For values of "usable" that include "14.6 seconds/token". It's a cool accomplishment! And newer hardware would speed it up some. But I think I'd want something a bit faster before declaring it usable in practice.