An agent collects a reward at every step. The return from a step is everything it will collect from that point to the end of the episode — but with later rewards counting for less, discounted by gamma once per step of delay:
Task: write discounted_returns(rewards, gamma) returning the return for every step, in the same order as the rewards, each rounded to 4 decimal places.
rewards[t] is the reward collected at step t. Rewards can be negative.gamma runs from 0 to 1. At gamma = 0 the agent is entirely short-sighted and each return is simply that step's own reward; at gamma = 1 nothing is discounted and every return is the plain sum of what's left.The efficient way is to walk the list backwards, carrying a running total:
Start with 0 and each step needs one multiply and one add — no powers of gamma anywhere, and one pass instead of the nested loop the first formula suggests.
That rewrite is worth internalising, because it's the same move behind TD learning, GAE and every other backward recursion in reinforcement learning: a sum stretching to the horizon becomes one step plus the thing you already computed.