← Back to context

Comment by glouwbug

6 hours ago

Learn which instructions SIMD nicely (sqrt / fabs, etc). Use ternaries in loops for masking. Use trig identities and lookup tables (don't recompute sin(3t) when you can use two vector multiples using a table of sin(t) eg. sin(t) * sin(t) * sin(t)). Use divisible constexpr constants in loops to eliminate the SIMD tail. Be careful with type casts and floats. `float x; x += 0.5` will introduce *cvt instructions even if the compiler statically knew better otherwise (use 0.5f). Compile with --fast-math and friends so errno doesn't invalidate your SIMD pipeline.

Most applications (including most applications that care about numerical performance) should not use -ffast-math.

That has a similar problem to the article, it's trying to fit far too much into too small a format.

What you've written mostly makes sense to someone who already has a solid understanding of SIMD and of C++ (although I can't say I follow all of it), but the target audience is people who don't. For them, each point needs a much lengthier explanation.

  • Likely the best tip would to `objdump -d` and inspect the assembly then checking performance counters. Prepending (__attribute__((used)) will allow you to inspect your functions.

    A quick restrict example:

        #define fn __attribute__((used))
    
        fn void copy1(int* to, const int* from, const int size)
        {
            for(int i = 0; i < size; i++)
                to[i] = from[i];
        }
    
        fn void copy2(int* to, const int* from)
        {   
            constexpr int size = 1024;
            for(int i = 0; i < size; i++) 
                to[i] = from[i];
        }
    
        fn void copy3(int* restrict to, const int* restrict from)
        {
            constexpr int size = 1024;
            for(int i = 0; i < size; i++) 
                to[i] = from[i];
        }
    
        gcc test.c -c -O3 && objdump -d ./test.o
    

    copy1 is 52 lines, copy2 is 28 lines, copy3 is 2 lines (just a call to memcpy).

    This is a good starting point for self teaching. The impact of your TLB, L1, and overall instruction count (with IPC) can further be measured with `./perf stat -d -d -d ./a.out`. If you want a quick rule of thumb, no instructions are fast instructions.