Comment by danielmarkbruce
2 days ago
It's not a prediction of the next move though, and that is the point. It's a prediction of what will happen if you make that move.
So, it's not a next move predictor. It's a game result predictor.
2 days ago
It's not a prediction of the next move though, and that is the point. It's a prediction of what will happen if you make that move.
So, it's not a next move predictor. It's a game result predictor.
Eh, no, that's not right. I might need to brush up on my Sutton & Barto but the RL task is traditionally defined as, informally, "given a current state observation predict the next action, state and reward". A policy is always predicting the next timestep's reward. Otherwise, how would it know what to do next?
Brush up :)
The policy is optimized to maximize the the total reward, defined as the sum of the reward at each step, discounted by some factor.
Alright, I'll have to check up on that. Thanks for being nice about it.
1 reply →