Comment by phire

2 years ago

Pretty sure it's Haswell and Zen 2. They both implement IT-TAGE based branch predictors.

I just assumed the M1 branch predictor would also be in the same class, but I guess not. In another comment (https://news.ycombinator.com/item?id=40952404), I did some tests to confirm that it was actually the threaded jumps responsible for the speedup.

I'm tempted to dig deeper, see what the M1's branch predator can and can't do.