Comment by parineum

7 hours ago

It could be that a model prefers the tribe mentioned in closest proximity to the word candidate most of the time. It could be that it prefers the one that's third in a series. It could be that it prefers the one with even numbers of letters.

The model is biased. That's it's entire function, to bias certain tokens over other tokens based on a bunch of vectors and context. There's no telling what is influencing that bias.

The models will be statistically more likely to choose one of the options for completely unknowable reasons.