← Back to context

Comment by jacekm

2 hours ago

How will this affect the newly build data centers? What effect do you think it will have on memory prices?

My uneducated guess says, not much. For running massive models you still need a ton of high-bandwidth interconnects between many individual chips/GPUs/etc since you need to do math across a few TB worth of weights. That's simply going to require more power (and more die area in I/O, and therefore more cost). Being able to run small models in tiny power envelopes is incredibly useful to people, but I believe it will be covering a different niche than what datacenters can provide. Likewise, you'll still need crazy amounts of high-end memory to populate whatever goes in these datacenters.

The only thing that will crash prices is reduced demand (duh) or, more interestingly, increased production. In particular, if CXMT is able to get their DR5 fabs up to a reasonably high yield, that could add some downward price pressure (as could government subsidies). As well, if Micron/Kingston/Hynix think that CXMT is going to start cutting into their market share, they might be willing to either increases supply or drop prices. Unfortunately CXMT looks to be taking quite a while to get their new fab up to max capacity so that may take a year+ before anything manifests.

If you're interested in following the (publicly available) info on these sorts of things, check out what companies like Axelera, DeepX, and MemoryX are doing today and have on their roadmaps, as well as the sorts of chips/SoCs Qualcomm, Kinara (now NXP), and Ambarella currently have announced (or have on the market). And remember, that pretty much all of these chips on the market today were in initial development more or less when ChatGPT first launched. If you knew what you knew today (or a year ago) about what requirements current- and next-generation models would have (from a silicon perspective), what might you do differently? Think for instance, host system interconnects, amount and speed of on-package or on-die memory, image/video decode capabilities, int8 vs fp8 vs fp16 vs bf16 compute units, etc. And, consider that most "AI" stuff in development a few years ago was all 15nm or 12nm - because who was gonna pay big money to get fab capacity at 3nm to run some object detection models? So most of the stuff on the market today is on very old nodes and therefore not super power efficient.