Any model ran takes compute. Everybody, datacenters and local need compute.
And I don't see a future where datacenters lose much to open models. I've ran local server rooms before clouds were a thing. Cooling and power is a very expensive pain in the ass. You're not dropping a rack of local models in most buildings without breaking the local power grid.
What about when the open models get to the point that you don’t need these fancy GPUs? It seems inevitable this will happen over time, we see cracks right now within even just GPT w/ expert streaming and flash attention. ASICs are coming soon too.
AWS and Azure would love to cut out OpenAI and Anthropic and sell inference on open models directly to businesses with no cut to share, just as AWS serves other open source software.
Open models are great for everyone except anthropic and OpenAI.
open models don’t imply local models… it just allows more people to run them, most likely with nvidia GPUs
Any model ran takes compute. Everybody, datacenters and local need compute.
And I don't see a future where datacenters lose much to open models. I've ran local server rooms before clouds were a thing. Cooling and power is a very expensive pain in the ass. You're not dropping a rack of local models in most buildings without breaking the local power grid.
What about when the open models get to the point that you don’t need these fancy GPUs? It seems inevitable this will happen over time, we see cracks right now within even just GPT w/ expert streaming and flash attention. ASICs are coming soon too.
Nvidia already have a plan in place for fast inference. They acquihire Groq last year
AWS and Azure would love to cut out OpenAI and Anthropic and sell inference on open models directly to businesses with no cut to share, just as AWS serves other open source software.
Open models are great for everyone except anthropic and OpenAI.