Comment by KerrAvon 3 hours ago We're on the cusp of Kimi K3 becoming usable on sub-10k hardware.https://github.com/gavamedia/deltafin 4 comments KerrAvon Reply edot 3 hours ago 14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp! dotancohen 8 minutes ago It's need it to be an order of magnitude cheaper, but for some queries ("What to discuss at tomorrow's meeting") I can wait 12+ hours. antonvs 7 minutes ago At that rate, it would take me a mere 15 days to generate the number of tokens I typically use in a day."Yes, I'm using AI to speed up development. I'll submit that PR in two weeks time!" ekidd 3 hours ago For values of "usable" that include "14.6 seconds/token". It's a cool accomplishment! And newer hardware would speed it up some. But I think I'd want something a bit faster before declaring it usable in practice.
edot 3 hours ago 14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp! dotancohen 8 minutes ago It's need it to be an order of magnitude cheaper, but for some queries ("What to discuss at tomorrow's meeting") I can wait 12+ hours. antonvs 7 minutes ago At that rate, it would take me a mere 15 days to generate the number of tokens I typically use in a day."Yes, I'm using AI to speed up development. I'll submit that PR in two weeks time!"
dotancohen 8 minutes ago It's need it to be an order of magnitude cheaper, but for some queries ("What to discuss at tomorrow's meeting") I can wait 12+ hours.
antonvs 7 minutes ago At that rate, it would take me a mere 15 days to generate the number of tokens I typically use in a day."Yes, I'm using AI to speed up development. I'll submit that PR in two weeks time!"
ekidd 3 hours ago For values of "usable" that include "14.6 seconds/token". It's a cool accomplishment! And newer hardware would speed it up some. But I think I'd want something a bit faster before declaring it usable in practice.
14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp!
It's need it to be an order of magnitude cheaper, but for some queries ("What to discuss at tomorrow's meeting") I can wait 12+ hours.
At that rate, it would take me a mere 15 days to generate the number of tokens I typically use in a day.
"Yes, I'm using AI to speed up development. I'll submit that PR in two weeks time!"
For values of "usable" that include "14.6 seconds/token". It's a cool accomplishment! And newer hardware would speed it up some. But I think I'd want something a bit faster before declaring it usable in practice.