Comment by joules77

4 days ago

At a basic level it generates a probability distribution of what the next token should be.

There are a zillion questions that can be asked where you can get a prob dist where multiple tokens have the same probability (flat probability distributions). Then it has to randomly pick one and you can get large variation.