← Back to context

Comment by sandworm101

10 hours ago

Those old LTT videos of high core-count threadrippers running GPU benchmarks become more relevant each day.

The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn).

But the memory bus speed is fully committed when generating tokens or thinking.