← Back to context

Comment by fsbonetto

6 hours ago

The bottleneck, for inference at least, is memory bandwidth. And that you can't make any faster by making it specific to your model.

So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.

Presumably though the kernel has a pretty specific set of operations done against the weights in memory. Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.

The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.

I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.

Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.

But such is the cycle

  • You're starting to hint at compute-in-memory as a general replacement for CPU/DIMM layouts. That could be useful for far more than LLMs, world models, or any sort of AI. It takes a bit of a different software development stack than a standard architecture though.

  • > Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.

    Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack.

    • I don’t think it would require a new training stack, and I’d imagine it makes more sense to distribute cores with memory. The cores can be simplified to the functions of the kernel since the inference kernel can be expressed as a reduced set of optimized functions in the pipeline rather than a general CUDA core. If the model is burned into ROM, the compute pipeline can be baked into the core.