← Back to context

Comment by Scene_Cast2

3 hours ago

What about numpy, numba, and torch.compile?

Those are manually optimized per arch, aren't they?

  • The libraries, yes, but the code you write is portable (at least until you get into squeezing the last few percent and switch to Triton / Helion in case of GPU, and even those are decently portable).

    There's also Halide, where you write the algo but the framework gets you the scheduling and SIMD.