← Back to context

Comment by vardump

8 hours ago

Benchmarking like that is often broken because of continuous CPU core clock speed adjustments, system interrupts, SMIs, etc.

I tried to fix it by switching hyperthreading off, playing with the scaling governor, boost, setting a CPU frequency to no avail. The jitter was too much and the results were not reproducible, so I just gave up.

Of course your mileage may vary; this was on an AMD Zen 3 CPU.

Intel published some guidance on doing precise benchmarks. Basically run your code inside the kernel, turn off interrupts, use the cpuid instruction to prevent out-of-order execution, use rdtscp instruction instead of rdtsc, etc.

  • > prevent out-of-order execution

    Won't this make the results inaccurate?

    • You just want to prevent the CPU from capturing the end time before it has finished the benchmarking code. Or prevent it from running benchmark code before it has captured the start time.

The way you work around this (apart from doing what you can to make the system as predictable as possible) is to accept that the data is noisy, and then working around it by doing multiple trials and following up with a statistical analysis on the results.

You can use the desired confidence to inform the warm-up, number of trials, benchmark duration, and so on.

If you end up with a multimodal distribution it can be worth tracking percentiles.

  • I did warm-up rounds, to prime caches and so on. I fiddled with statistics: IIRC simple average seemed to give the best results overall. I thought keeping the lowest n results would be the best, but that turned out to be wrong.

    The results were indeed multimodal.

    All I wanted was to have a reproducible benchmark to see how the code changes affected performance over time.

Used to do this sort of thing for computations that needed to run in the 10 microsecond range (HFT stuff), circa 2008. Had very predictable results because:

a) language was not garbage collected (C++)

b) we avoided heap lock contentions in critical paths by pre-allocating object pools at startup

c) I/O operations were offloaded to separate threads, connected by mutex locked linked lists

d) processing thread was bound to its own CPU core

That's about as deterministic as we could get.

  • Did all of those trying to reduce the jitter. I think it was about the CPU clock not being stable.

    Additionally I avoided core 0, because it was the noisiest and did some cgroups core pinning for the test workload.

    There were zero page faults during the test runs and the CPU core was uncontested by other threads.

    I think in 2008 CPUs were not so crazy about power and heat management.