← Back to context

Comment by wat10000

9 hours ago

x86-64 and ARM64, at least, let you test an arbitrary bit in a register with one instruction.

The fact that bit testing also needs one instruction does not mean that it is equally efficient.

On x86-64, there are 2 ways to test the value of a bit. If you use the bit testing instruction (BT), that instruction is both longer and slower than testing if a register or memory value is null.

If you use the test-under-mask instruction (TEST), this is fast, but the instruction is significantly longer (by including an immediate constant for the mask). Longer instructions can also cause lower speeds, when various bottlenecks are encountered, e.g. the maximum number of bytes fetched per clock cycle or the capacity of the instruction cache or of the micro-operation cache.

Moreover, testing whether a value is null frequently requires zero instructions, not one instruction, because if the value is the result of computing some expression then the flags register already stores if the value is null or not (and its sign).

On ARM Aarch64, the instruction that tests a bit in a register has a much smaller jumping range than the one that tests whether the whole register is null, so testing a bit in a register may require the insertion of an extra jump instruction to reach the target where execution should continue.

Once upon a time testing whether all bits of a number are zero was slower than checking a single sign bit. Even when MIPS was originally designed, Hennessy and his team had some trouble with making BEQZ/BNEZ fast enough for their intended pipeline.

  • This happens because testing the sign bit needs just a wire from that bit to the flags, while testing if a register is zero requires a wide OR gate with as many inputs as there are bits.

    In CMOS you cannot have an OR gate so wide, so it must be synthesized from a cascade of narrower gates, which add several levels of delays.

    While in modern CPU technologies the speed of generating a zero flag is not a problem, when designing with FPGAs, which are much slower, it is useful to be aware that testing for the sign is cheaper than testing for a wide zero.

    Many CPUs have an instruction for implementing loops like decrement-and-jump-if-not-zero (which is LOOP in x86-64). When implementing a simple CPU in an FPGA it is cheaper and faster to replace that instruction with 2 instructions for loops like increment-and-jump-if-negative and decrement-and-jump-if-not-negative (it is good to have both these instructions to be able to access an array both in forward order and in reverse order, while using the loop counter also as index register).

    The same applies when making a counter in FPGAs, it can count at higher frequencies if you test for the sign bit to determine the end of the counting, instead of testing when the count reaches zero.