Comment by danielmarkbruce

2 days ago

It's not a prediction of the next move though, and that is the point. It's a prediction of what will happen if you make that move.

So, it's not a next move predictor. It's a game result predictor.

Eh, no, that's not right. I might need to brush up on my Sutton & Barto but the RL task is traditionally defined as, informally, "given a current state observation predict the next action, state and reward". A policy is always predicting the next timestep's reward. Otherwise, how would it know what to do next?