The libraries, yes, but the code you write is portable (at least until you get into squeezing the last few percent and switch to Triton / Helion in case of GPU, and even those are decently portable).
There's also Halide, where you write the algo but the framework gets you the scheduling and SIMD.
Those are manually optimized per arch, aren't they?
The libraries, yes, but the code you write is portable (at least until you get into squeezing the last few percent and switch to Triton / Helion in case of GPU, and even those are decently portable).
There's also Halide, where you write the algo but the framework gets you the scheduling and SIMD.