← Back to context

Comment by spacedcowboy

7 hours ago

The benchmarks for 'xc' - "my" compiler for heterogenous computing (gpu and cpu, all in the same language, compiler managing data-hazards between them) are all on the order of a second to a few seconds, for the long pole.

So if I'm comparing against (say) C++, Swift, ObjC, on a given (arm64 or x86_64) architecture, I want to see that sort of timescale, a bit less is fine, a bit more is fine.

Of course, when you (today [grin]) get auto-vectorisation of 2D matrix multiplies, and you have an SME/SME2 target on arm64 that most compilers don't pick up so you're 150x faster than clang/g++, you might have to run it a bit longer so you can get reasonable comparison numbers :)

1: https://compile-xc.org/compiler/performance/