← Back to context

Comment by rerdavies

2 years ago

It would be interesting to see what happens if you turn on profile-guided optimization. This seems like a very good application for PGO.

The 518ms profile result for the branch-table load is very peculiar. To my eyes, it suggests that the branch-table load is is incurring cache misses at a furious rate. But I can't honestly think of why that would be. On a Pi-4 everything should fit in L2 comfortably. Are you using a low-performance ARM processor like an A53?