"This action earned a reward of 10" means nothing on its own. Was 10 good? If the agent expected 12 from that state, it was a disappointment. The advantage answers the only question that matters for learning: did this turn out better or worse than expected?
The one-step version compares what actually happened against the critic's prediction for the state you were in:
Task: write advantages(rewards, values, gamma) returning one advantage per step, each rounded to 4 decimal places.
rewards[t] is the reward at step t; values[t] is the critic's estimate for the state at step t. Both lists are the same length.values[t+1] for it. Treat the value after the final state as 0 — a terminal state has no future worth anything.This is why policy-gradient methods keep a critic at all. Multiplying the gradient by a raw return pushes up every action that preceded any positive reward, including the bad ones in a lucky episode. Multiplying by the advantage pushes up only the actions that beat the baseline — far less noise, far faster learning.