Comment by BoredomIsFun
2 hours ago
> not 1 token on 1 session like local models.
Local models can absolutely run in batch, what are even talking about?
> If you had only batch inference and enough of it to fill the compute to 80% then you get cheaper local models.
Even if you ran sequentally, single session, a _finetuned_ tiny (8B) local model on narrow tasks would abolutely mog SOTAs, any of it - Fable, Opus, Sol you name it.
No comments yet
Contribute on Hacker News ↗