Comment by johnnyApplePRNG

4 hours ago

Their strategy is to prevent open models from proliferating, so their massive investments in these AI frauds are not completely unwound.

That's my take, at least.

Nemotron is TERRIBLE, and purposefully so. It must be.

They cannot be THAT BAD at training AI models. I don't believe it.

I think almost the exact opposite. They want open models. Without credible open models, they only have a few customers, and those customers have leverage against nvidia. With open models, they have tons of customers, and nvidia has all the leverage.

Smaller customers are also less able to develop their own hardware and threaten NVidia's business.

  • > Without credible open models, they only have a few customers, and those customers have leverage against nvidia. With open models, they have tons of customers, and nvidia has all the leverage.

    I wonder if Nvidia is also trying to cover themselves against a market crash the bankrupts the AI labs but leaves the tech standing? (Similar to the dotcom crash or the big railroad crash back in the day.) If Anthropic and OpenAI struggle, Nvidia can always sell local inference hardware. But local inference hardware often has a lower utilization, so you need more GPUs for the same number of tokens.

    • I think that's their exact plan. Cover all bases. The bubble can burst but that really only affects the UI model giants, the demand for AI won't drop even if frontier model companies won't find profitability enough to pay the bill

  • Sure but it also means these companies nvidia buys don’t work with AMD and other competitors anymore. There are myriad motivations and it’s not all one or the other but this is a classic component of Silicon Valley acquisition strategies.

  • They want LEVERAGE.

    They can't ride the hyperscaler gravy train forever; at some point between Google, AMD, and Apple NVIDIA is going to lose its monopoly on serving large customers.

    At that point, it would be useful if a few open models existed which were only a couple of months behind the frontier.

    But it's very important that the open models never be TOO good, because the AI companies are buying compute on the assumption that their software will add value. If it becomes a commodity business with frontier open models, NVIDIA won't be able to get away with such a crazy markup.

    • > If it becomes a commodity business with frontier open models, NVIDIA won't be able to get away with such a crazy markup.

      Of course they would. You need hardware to run the model. nVidia sets the floor.

Have you used Nemotron-3.5-lightening? I don’t use it as much as Poolside’s (excellent!!) Laguna XS 2.1 6bit, but the new Nemotron model is good.

I think NVIDIA does want small open models running on-prem to explode as a market! Lots of smaller GPU installations for companies who wisely want on-prem inference.

Of course NVIDIA will also keep making a ton of money selling to hyper scalers, but not forever: Chinese chips are getting better, Google, Microsoft, Amazon, etc. designing their own inference chips.

NVIDIA is handling this brilliantly.

I think this is a classic [Commoditize Your Complement](https://news.ycombinator.com/item?id=17047348) - Nvidia wants open models because their business is hardware and it's complement is AI models, so they want AI models to be commoditized so that hardware is the industry with leverage. OpenAI/Anthropic/etc want closed models so that the AI model development/data has leverage over the hardware providers.

Nvidia is trying to increase its customer base. Look at the Mag 7, Amazon, Google, Microsoft, and even Meta are all working on their own inference chips. I don't think they can completely ditch Nvidia for LLM training, but they can make their own chips for inference. I believe that's also why OpenAI made its own.

No one wants to pay the Nvidia tax

  • If you make your own chip don't you need to make your own software stack too. Which is what I thought kept everyone using Nvidia.

I believe NVIDIA's long term strategy will be to pivot from the data center to the public, and the public will utilize open models on NVIDIA hardware at home. This will come after the RAMpocalypse completes (when the new fabrication plants (China, Tesla/SpaceXAI/Intel) fully ramp up and start selling their RAM for cheap in the next few years). Data centers will be for training mostly.

  • There’s a lot about this that would make sense.

    But, not really at the current technologies. Kimi and GLM are fucking awesome, but I don’t have 3TB of VRAM to run them, and I don’t expect to even when ram prices drop.

    So now you’re back to the scaling issue before talking about power and compute distribution.

    • Do you think you'll realistically need 3 TB RAM to run a sufficiently good model 1 to 2 years from now? I certainly don't. Considering what can already be done with 128 GB of relatively slow unified memory, imagine if efficiency improvements continue apace, the memory becomes 128 GB of HBM, the flash device becomes capable of sequential throughput matching today's DDR5, and such a system was affordable as a routine purchase for the average person.

  • Why do you think the public wants to self host models over using a cheaper solution hosted in the cloud?

    • Almost all the non-tech people I know acknowledge concerns about privacy but have no realistic alternative available to them. I think they won't want to self host until and unless it's easy to deploy, there's nothing to administrate, and the hardware cost is in the same ballpark as a new laptop was prior to the apocalypse.

I doubt Nvidia wants to be fully dependent on the success of two highly unprofitable companies that could implode at any moment. It makes much more sense for them to commoditize LLMs so that their target market grows to every mid-sized or larger company.

Maybe as an LLM it is, but I constantly use their streaming ASR model (called Nemotron Streaming) through Handy and it works wonderfully well.

The other way. Nvidia would love open source jevon paradoxed ai - that would run inference on their chips.

  • But their goal would be to ensure it only runs on their chips, and not any competitors. I can't see how they could do that if the best models truely were "Open".

    I can see one of Nvidia's biggest fears is the inference hardware becoming commoditised.

    • The existence of inference ASICs has been priced-in to Nvidia's stock since Google started designing TPUs. It's not that new, really.

      What Nvidia has the market cornered on is flexible GPU architectures. Nobody else has stepped up to the plate on that, and it's how Nvidia will butter their bread with robotics and future model training efforts.

  • can someone weigh in on this. what's the actual play here

    are they actually suppressing the western open models?

    china doesn't give a fuck either way

    imo, they see the weakness emerging at the intersection of all the labs, everybody knew there was no moat, so they're gonna control its direction and basically tell the Jev guys what they want them to work on

    • It’s just an illogical conspiracy theory. Open weights models still have to run on someone’s chips. NVIDIA is model-agnostic.

      They are just trying to grow the pie because they have nobody else competing for slices.

      My guess is they want as many frontier models using their chips as possible. The only threat to their business is companies making their own chips which Google does and the others are working toward. The last thing they want is only 3 frontier labs who are all not buying NVIDIA.

      2 replies →

I think the idea that they’d purposely spend company time and resources making a bad model is an extraordinary claim, requiring extraordinary evidence. The more likely explanation is that they aren’t willing to distill from their own customers, so they are at a disadvantage.

Insane conspiracy theory. There's no incentive whatsoever for Nvidia to release weak models. If you bothered to pay any attention, they are aggressively trying to catch up. Whether they succeed, that's of course a separate question.

  • why would Nvidia try to compete against their biggest customers - openai and anthropic ?

    • Dog fooding their own product. This isn’t new for Nvidia. For a long time they have been selling GPU cards, and also licensing the chips for other manufacturers to make their own cards. Similar for their automotive products, Shield, and DGX Spark.

    • the power that giveth can also taketh away.

      similar to how eventually AWS started making their own chips for data centers, and Apple did that for their hardware, it's not a ridiculous thing to plan for the AI companies to start making their own chips to optimize for their use cases and cut out the middleman for margins.

    • So that nvidia gets bargaining power...? Nvidia needs to diversify its customer base