Comment by someonebaggy

7 hours ago

> Such as whether a callee-saved register really does need to be saved in some particular routine.

Ironically this is your preconceived abstraction of how a compiler has to operate. An ideal compiler could allocate registers differently for each called function: F1()->F2()->F3(), F1 uses r0-5, F2 uses r6-10, F3 uses r11-15, no register saving required in the whole chain. There's no need for a fixed ABI. Such a compiler would look very different from today's ones.

Modifying the ABI of a function requires being able to track down all of the call-sites of the function, which is less trivial than you might assume. ABI concerns also tend to baked in relatively early in the optimization pipeline because you just simply can't get the ABI wrong, and I can think of several instances where the ABI decision causes missed optimizations.

There is also the other issue that a good algorithm for optimizing a problem like register allocation tends to be super-linear (e.g., quadratic), and if you shift the model from "allocate on a per-function basis" to "allocate all functions", the N in the O(N²) goes from "size of function" to "size of program," which is now suddenly a lot more compiler time spent for very modest gains. If register spilling across a function call is a noticeable component of runtime, then you're probably better off inlining that function in the first place!

  • One of the hardest parts of optimizations is not being penny wise and pound foolish. This sort of thing is a great example of that.

    Additionally, a hard part is that all of this can change over time with new hardware! Some patterns that were crucial before everything gained branch predictors are irrelevant now, etc.

    • > penny wise and pound foolish

      I love this saying. The general problem, optimizing the wrong metric, shows up all over the place.

Yeah, that's what MSVC and GCC did on x86, called "custom calling conventions" on MSVC and the regparm attribute on GCC.

All of these were dropped on x64, on x64 (and ARM) you get standard calling conventions for just about everything with proper unwind tables for functions.

It doesn't really "cheat" on the registers unless it inlines a function entirely. LLVM has support for custom calling conventions and pragmas to specify them, this is used by GHC on Haskell and other things, but it's practically unheard of in "normal" C/C++ code.

If it were practical and if there were significant performance benefits (on modern heavyweight CPUs) to tailoring the calling convention on a per-function basis, I'd expect to see it done in optimising JIT engines like Java HotSpot. As far as I know, they don't bother.

I'm no expert but I suspect jcranmer's comment has it right that you end up doing cross-function register-allocation while foregoing the other benefits of just inlining. I also suspect the payoff would be minimal on modern heavyweight hardware. I can see it making more of a difference on a very minimal embedded processor, or if optimising for the smallest binary possible.