← Back to context

Comment by Panzerschrek

9 hours ago

-1 is not -1 for different sizes of integers. Which exact -1 do you want? 32-bit? 64-bit? int-sized (as int is defined in current implementation)? With current approach it's simple, boolean type is expected to have 1-bit size (which is rounded to 8 bits if stored in memory).

If bool was a signed 1-bit type (i.e. having possible values of either 0 or -1), it could be extended to any size, same as an integer?

Of course we can't have that now, because unsigned bool is baked into the C language standard, and even into processors (x86 SETcc instruction). While "char" being signed or not is still implementation-defined IIRC...

You have some expectations of the size of the bool (it's simply 1 bit but also 8 bits.)

What expectations do you have of the value?

  • Bool should be logically 1-bit, when stored in memory only least significant bit should be used and the rest is allowed to be garbage. Such approach gives compilers as much room for optimizations as possible. Forcing them writing some specific bit-pattern may lead to suboptimal code generation.

    • Bool is 1-bit, but that bit can be defined as signed or as unsigned.

      If bool is defined as unsigned, casting it to any size of integers will give 0 for false and 1 for true (using the standard zero-extension operation that converts smaller unsigned integers to bigger unsigned integers).

      If bool is defined as signed, casting it to any size of integers will give 0 for false and -1 for true (i.e. an all-1 bit pattern) (using the standard sign-extension operation that converts smaller signed integers to bigger signed integers).

      Defining bool to ignore the other bits except the LSB leads to a lower performance on most processors, because in almost all instruction sets it is more efficient to test whether an integer is null or non-null, than to test the value of a bit.

      The only efficient way to use a single bit and to ignore the others would be to store the boolean in the most-significant bit, i.e. in the sign bit of a signed integer, because testing the sign is normally as simple as testing whether a value is null. If this convention were used, a boolean result could be 0 for false and -1 for true, but in input arguments negative would be true and non-negative would be false.

      8 replies →

    • That's how gcc does it (at least in this one case), but as TFA points out, this is not standard-conforming.

    • I vaguely remember that Clang and GCC used to disagree what the contents of the upper parts of x64 registers when returning some integer types should be (zeroes or garabge), because the PDF that defined Sys V ABI on x64 left such irrelevant details out, so linking together objects produced by those compilers, both of which claimed to follow the same ABI, would produce malfunctioning executable.

      > Forcing them writing some specific bit-pattern may lead to suboptimal code generation.

      So? Forcing them to compile "return 42;" as "mov eax, 42; ret" also leads to suboptimal code generation: a plain "ret", returning whatever is in rax already, is optimal. It doesn't generate the specific bit pattern for 42 but that's a small price for the improved efficiency, isn't it?

    • The in-memory representation is a completely different question. The language could easily say that true has an integer value of -1 while still storing it as a single bit.