← Back to context

Comment by barbacoa

7 hours ago

They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs.

So full CPU local AI inference may become viable option in coming years.

the GPU competition is using 16 gpus, so the actual comparison is that the CPU has <1/10th the bandwidth

This is essentially guaranteed. There are lots of useful smaller models that we should be able to run locally. Over time they'll be more and more capable and require less API usage.

  • Im wondering if we are finally seeing the end of the "hard disk" era, and are entering a new era of vast instant on systems.