An agent keeps an estimate of how good each state is: values[s], the total future reward it expects from there. Temporal-difference learning improves one of those numbers after every single transition, without waiting for the episode to end.
The agent was in state, received reward, and landed in next_state. Its own two estimates now disagree:
values[state],reward + gamma * values[next_state].That gap is the TD error, and the agent nudges its old estimate a fraction of the way toward the new one:
Task: write td_update(values, state, reward, next_state, alpha, gamma) returning the full updated list of values, each rounded to 4 decimal places.
values is indexed by state number. Return a new list — the same length, with every other state untouched.state.alpha is the step size (how much of the gap to close) and gamma discounts future reward.state and next_state can be the same, in which case the entry you're reading is the entry you're writing. Read it before you overwrite it.What's strange and powerful here is that the target is built partly from the agent's own current guess about next_state. It is learning from a number it made up — bootstrapping — and it still converges, because the real reward in every update drips a little bit of ground truth into the estimates.