← Back to context

Comment by marginalia_nu

4 hours ago

The way you work around this (apart from doing what you can to make the system as predictable as possible) is to accept that the data is noisy, and then working around it by doing multiple trials and following up with a statistical analysis on the results.

You can use the desired confidence to inform the warm-up, number of trials, benchmark duration, and so on.

If you end up with a multimodal distribution it can be worth tracking percentiles.

I did warm-up rounds, to prime caches and so on. I fiddled with statistics: IIRC simple average seemed to give the best results overall. I thought keeping the lowest n results would be the best, but that turned out to be wrong.

The results were indeed multimodal.

All I wanted was to have a reproducible benchmark to see how the code changes affected performance over time.

  • I've been doing some benchmarking recently and I've observed that on certain workloads, my CPU cores are not equivalent. (And this is an Intel CPU from 5 years ago, all cores are the same on paper.) In a particular very bad case I have a reproducible 50% performance difference between (physical) core 0 and core 5.

    You might want to taskset to particular cores too, especially if you are looking at computational code with many almost-but-not-quite-random memory accesses.