← Back to context

Comment by BoredomIsFun

2 hours ago

> not 1 token on 1 session like local models.

Local models can absolutely run in batch, what are even talking about?

> If you had only batch inference and enough of it to fill the compute to 80% then you get cheaper local models.

Even if you ran sequentally, single session, a _finetuned_ tiny (8B) local model on narrow tasks would abolutely mog SOTAs, any of it - Fable, Opus, Sol you name it.