Comment by spankalee
6 hours ago
I would rather say: "benchmark with confidence intervals" or maybe "benchmark with comparisons".
It's very hard to say much about an absolute number. You need to compare against some alternative or control, and because of CPU load, throttling, GC, and a thousand other variables, you really should be comparing against that control _in the same run_, and importantly round robin across multiple runs to spread out the noise fairly across each implementation.
Then once you have a bunch of measurements you have a distribution and shouldn't just take a mean to compare, but should calculate something like the 95% confidence interval. If you see that those confidence intervals overlap, the you might not really know which is faster. If they don't overlap, then you probably do know which is faster.
If you have a good benchmark, then running it more times can narrow the confidence intervals and let you tease out very small improvements at the cost of longer runs. If confidence intervals don't narrow, then you hit the limits of signal-to-noise.
This is the only way I've been able to get reliable, actionable benchmarks outside of a very, very controlled hardware lab. It's what Google's Tachometer benchmark runner does, and I wish more runners did this: https://github.com/google/tachometer
Instead of giving users a confidence interval, ask the user to specify a p value. Then run something like Welch’s t test.
Credible interval of difference of means, please.
https://link.springer.com/article/10.3758/s13423-015-0947-8
https://web.archive.org/web/20250418135701/https://jkkweb.si...
mine is: "benchmark in load and wall-insensitive witnesses"
it's tempting, but otherwise you're just attesting the limits of the current environment