TD learning updates its estimates from its own guesses. Monte Carlo evaluation refuses to guess: it plays whole episodes to the end, measures what actually happened, and averages.
Task: write mc_evaluate(episodes, gamma) returning a dictionary mapping each state to its estimated value, rounded to 4 decimal places.
episodes is a list of episodes. Each episode is a list of [state, reward] steps, in the order they happened.The neat way to get the returns is to walk the episode backwards with a running total: g = reward + gamma * g, starting from g = 0. One pass, no powers of gamma to track.
First-visit is the rule that makes the averaging statistically clean — counting both visits of a repeated state would give that episode two correlated samples pretending to be independent ones, and bias the estimate.