A uniform replay buffer wastes most of its samples on transitions the agent already understands. Prioritized replay samples the surprising ones more often — but sampling unevenly biases what the network learns, so it then has to correct for its own bias. Both halves are pure arithmetic.
Sampling probabilities. Raise each priority to the power alpha and normalise:
alpha controls how aggressive the prioritisation is. At alpha = 0 every priority becomes 1 and you're back to uniform sampling.
Importance-sampling weights. A transition drawn too often must count for less:
where N is how many transitions are in the buffer. Then divide every weight by the largest one, so the biggest weight is exactly 1 and no update is ever scaled up.
Task: write priority_weights(priorities, alpha, beta) returning [probabilities, weights] — two lists, each rounded to 4 decimal places.
priorities holds positive numbers, one per transition in the buffer.Note how the two halves pull against each other by design. A high-priority transition gets a large P(i), which gives it a small weight — it's seen often, so each sighting counts for less. beta sets how completely that correction is applied, and in practice it's annealed from small to 1 over training, so early learning gets the speed of biased sampling and later learning gets the correctness.