A deep RL agent doesn't learn from each transition as it happens — consecutive steps are far too similar, and training on them in order makes the network chase whatever it's doing right now. Instead it stores transitions in a replay buffer and trains on scattered batches pulled back out of it.
Training code wants that batch grouped by field, not by transition: all the states together, all the actions together, and so on, because that's the shape a network consumes.
Task: write sample_batch(buffer, indices) returning [states, actions, rewards, next_states].
buffer is a list of transitions, each one [state, action, reward, next_state].indices says which transitions to pull, in that order — the sampling has already been done for you.indices: all the states first, then all the actions, then the rewards, then the next states.The reshape is the whole job: you're handed data organised one transition at a time, and you have to hand back data organised one field at a time.