Comment by vardump
5 hours ago
I did warm-up rounds, to prime caches and so on. I fiddled with statistics: IIRC simple average seemed to give the best results overall. I thought keeping the lowest n results would be the best, but that turned out to be wrong.
The results were indeed multimodal.
All I wanted was to have a reproducible benchmark to see how the code changes affected performance over time.
I've been doing some benchmarking recently and I've observed that on certain workloads, my CPU cores are not equivalent. (And this is an Intel CPU from 5 years ago, all cores are the same on paper.) In a particular very bad case I have a reproducible 50% performance difference between (physical) core 0 and core 5.
You might want to taskset to particular cores too, especially if you are looking at computational code with many almost-but-not-quite-random memory accesses.