A policy network outputs probabilities over actions, and you want to make the good actions more likely. But "good" isn't a label you can train against — there's no correct action to compare with. The policy gradient trick is to build a loss whose gradient does the right thing anyway:
Task: write pg_loss(log_probs, advantages) returning that loss, rounded to 4 decimal places.
log_probs[t] is the log-probability the policy assigned to the action it actually took at step t. These are negative numbers, since probabilities are below 1.advantages[t] is how much better than expected that step turned out. Can be positive or negative.N steps, and negate.Why the minus sign: optimisers minimise. Pushing the loss down pushes up — and for a positive advantage, the only way to raise that product is to raise , which means making that action more likely. For a negative advantage the arithmetic flips and the same loss makes the action less likely. One formula, both behaviours, no labels anywhere.
And notice the loss value itself is nearly meaningless — it isn't an error you're driving to zero, and it can be negative. Only its gradient matters, which is why nobody plots it to judge whether training is going well.