Slacker News Slacker News logo featuring a lazy sloth with a folded newspaper hat
  • top
  • new
  • show
  • ask
  • jobs
Library

Comment by ChildOfChaos

3 days ago

Not as simple as that. Everyone would happily use local, but the issue is local sucks.

1 comment

ChildOfChaos

Reply

nekusar  3 days ago

https://github.com/brontoguana/krasis

On my desktop RTX 5060 TI (16GB) and 96GB ram, I routinely get 25-30 tokens/sec using an 80B model quantized to int8. Uses 65GB system ram and 15GB gfx ram.

And its plenty fast for many of my purposes.

I could easily run a 30B model bf16 (full) and do like 50tok/s

Slacker News

Product

  • API Reference
  • Hacker News RSS
  • Source on GitHub

Community

  • Support Ukraine
  • Equal Justice Initiative
  • GiveWell Charities