As long as you don't have "realtime" workloads, owning the GPUs quickly becomes the economical option. The main cost problems is in e.g. chat applications where the workload is spikey, and users expect an near-instant response, for which you need to scale the GPUs to the highest spikes of the workload.
You really don't need to be large. $100k can buy you a lot of compute and it's less than hiring an engineer. With that kind of money you can build an LLM server for a dozen people.
One engineer's salary to accelerate a team of twelve is so cheap you can't afford not to.
Open models on-prem is the future, not a single doubt in my mind.
I'm old enough to remember when my company had everything on-prem (both analytical and operational databases and servers) due to cost and security concerns. Nowadays we have everything on GCP.
The biggest problem we had with on-prem was maintenance as it took a lot of staff and time to ensure decent reliability.
As long as you don't have "realtime" workloads, owning the GPUs quickly becomes the economical option. The main cost problems is in e.g. chat applications where the workload is spikey, and users expect an near-instant response, for which you need to scale the GPUs to the highest spikes of the workload.
You really don't need to be large. $100k can buy you a lot of compute and it's less than hiring an engineer. With that kind of money you can build an LLM server for a dozen people.
One engineer's salary to accelerate a team of twelve is so cheap you can't afford not to.
Open models on-prem is the future, not a single doubt in my mind.
I'm old enough to remember when my company had everything on-prem (both analytical and operational databases and servers) due to cost and security concerns. Nowadays we have everything on GCP.
The biggest problem we had with on-prem was maintenance as it took a lot of staff and time to ensure decent reliability.