← Back to context

Comment by vlovich123

8 hours ago

Really depends on what you benchmark and how reliable you want the measurement and what domain you are benchmarking. For example, criterion can sometimes spend quite a bit of time because it needs to stabilize the measurements. The author’s claim is “well you don’t need that accuracy” but I’ve seen wasted time chasing ghosts or making claims on performance improvements that were either neutral or net negative due to this noise.

> Anything faster than, say, 10ms risks being skewed by fixed costs (e.g, interpreter startup).

Sounds like the author’s experience is strictly in Python. For example with Java you have to make sure the JIT has sufficient optimized your program.

Additionally there’s plenty of situations where it can take a really long time to generate a representative dataset worth benchmarking and it can take time to evaluate the performance (eg databases). Short and quick microbenchmarks can be useful as building points, but at some point you need to evaluate steady state performance of the full thing. Other domains this comes up with is game rendering performance where a 300ms sample tells you nothing about whether you have frame drops after minute 25 or have a memory leak.

Some optimizations like hoisting still improve code readability even if the benchmark are inconclusive. And of course how a piece of code behaves in vivo and under test can vary quite a bit in both directions. Particularly with space/time tradeoffs, where you reduce or increase cache pressure with code running concurrently to your code.

A lot of my meditations on optimization date back to a profiler telling me that a redundant function call was responsible for 5% of the run time of a task. After removing it, run time decreased by 20%. Then I had to think about all the ways in which profilers can lie. It’s still a black art after all this time.

The tools tell you whether it might be worthwhile to look at a problem, but keeping your work is a completely different matter entirely. Unfortunately some people get Sunk Cost Fallacy, or worry about losing face, so once committed to a course will see it merged into the codebase whether it does anything or not. And they will push harder if they win the lottery and one test run says theirs is much faster. Nevermind that the next ten runs show the opposite.

  • > how a piece of code behaves in vivid and under test can vary quite a bit

    I've never seen "in vivid" used this way. Are you thinking of "in vivo" which is from Latin meaning "in life" or "in living" distinguished against Latin "in vitro" meaning "in glass" referring to the glass petri dishes or beakers used to do science experiments in a laboratory?

    • That’s because Apple autocorrect is a long con to prove that humans are too stupid to be left in charge. We can’t even english good. I assure you I did not type “in vivid”.

In high volume systems, 10ms is kind of crazy. I’ve run systems with operation metrics in the ms scale and the server side latency was lower than 1ms (computing business logic, or heavily cached data). Client side was closer to 3-5ms. As you mentioned, this was Java.

There also needs to be care taken in how these measurements are aggregated. Averages will almost always tell you nothing. High percentiles (95%, 99%, 99.9%) under load may show you something completely different than the average or even median case.

  • You can still benchmark say 100 requests and time that. For JVM for example you'll have to consider the JIT effects and GC and stuff that only happens later. So 100x 1ms requests is still a very small benchmark that's probably unreliable