← Back to context

Comment by margalabargala

3 hours ago

The same person managing the company's email accounts and whatnot.

I didn't say it's fire-and-forget. I'm saying all that is maybe a day of work every 3 months.

It doesn’t really work like that.

The companies which have the will and the budget to host their own LLMs typically require a ton of other, much smaller, models as well. They have internal security requirements, guardrails, audit, critical workflows start depending on your onprem setup, downtime is now something that’s not even allowed. There’s going to be a zoo of tooling, lots of bespoke work with internal clients who have no clue about docker, but now want their vibe-coded app to access a model they downloaded yesterday, and this model better be served and monitored 24/7 because now C-levels use it.

No one is going to budget several millions to then look at an email admin who hears ‘cuda’ for the first time in their life and ask them to just support the entire thing somehow.

And god forbid it’s an AMD setup.

This obviously isn’t relevant for a 10-person startup and their second-hand xeon with a single H100.

From my experience only the companies which are REALLY interested in privacy and data security bother with hosting their own models