← Back to context

Comment by linsomniac

4 hours ago

One of the smartest guys I know said: The longer you let a benchmark run, the more certain you can be that it captures other unrelated activity.

His theory was you want to run a benchmark several times, for short periods of time, and take the fastest run as the benchmark.

I ended up doing a lot of benchmarking for the Python Need For Speed Sprint, and that advice seemed to work out well for us there.

This is basically the philosophy of https://github.com/c-blake/bu/blob/main/doc/tim.md , but it uses a somewhat fancier estimator of the minimum than the sample min of dt's (which in some sense is guaranteed to be "a little" high).

There are other benchmarking points in that document about measurement footguns that can arise from targeting a specific duration like the 200..400ms of TFA with an easy but careless problem scale-up (like the Ben Hoyt example of footnote 4).

With that tool, in user space with just CPU freq pinning, I routinely see CPU bound times that are stable to single digit microseconds month to month on the same machine and see 0.01% to 0.2% effects on 10ms scale activity (yes, 1..20 bps) and the reported uncertainty usually captures the variation all right, but the distribution is not really Gaussian/Normal and instead has more shape parameters and/or you would want a 95% CI or some such as mentioned elsethread here.