There are two ways to estimate an advantage, and both are flawed. The one-step version uses a single real reward and then trusts the critic — low noise, but it inherits every bias in the critic's estimate. The full-return version uses nothing but real rewards — unbiased, but enormously noisy over a long episode.
GAE blends across all of them at once with a single knob, lambda.
Start with the one-step TD error at every step:
Then GAE is an exponentially-weighted sum of the errors from there to the end — which collapses, like discounted returns, into one backward pass:
Task: write gae(rewards, values, gamma, lam) returning one advantage per step, each rounded to 4 decimal places.
values is one entry longer than rewards: it holds the value of every visited state plus the bootstrap value of the state after the last step. So values[t + 1] is always safe to read.0.lam = 0 leaves you with exactly the one-step advantages, δ_t. lam = 1 gives the full discounted-return advantage. Anything between trades bias against variance.The two test cases at the extremes are the ones to study: at lam = 0 each step depends only on its own reward and the critic, and at lam = 1 information from the end of the episode propagates all the way back to the first step.