← Back to context

Comment by BoredomIsFun

3 hours ago

> not 1 token on 1 session like local models.

Local models can absolutely run in batch, what are even talking about?

> If you had only batch inference and enough of it to fill the compute to 80% then you get cheaper local models.

Even if you ran sequentally, single session, a _finetuned_ tiny (8B) local model on narrow tasks would abolutely mog SOTAs, any of it - Fable, Opus, Sol you name it.

> Local models can absolutely run in batch, what are even talking about?

I think the point was that if you aren't running your local machine at 100% for 24 hours a day then a cloud - with multiple clients - that is, will be more efficient.