Comment by BoredomIsFun
3 hours ago
> not 1 token on 1 session like local models.
Local models can absolutely run in batch, what are even talking about?
> If you had only batch inference and enough of it to fill the compute to 80% then you get cheaper local models.
Even if you ran sequentally, single session, a _finetuned_ tiny (8B) local model on narrow tasks would abolutely mog SOTAs, any of it - Fable, Opus, Sol you name it.
> Local models can absolutely run in batch, what are even talking about?
I think the point was that if you aren't running your local machine at 100% for 24 hours a day then a cloud - with multiple clients - that is, will be more efficient.