A hand-written backward pass is easy to get subtly wrong. A factor of goes missing, a ReLU gate is forgotten, a table is read the wrong way round. None of these crash. Training just goes quietly worse, and nothing tells you why.
The standard defence is to audit backprop with the crude method it replaced: nudge one weight and watch what the loss does. That costs two loss evaluations per parameter, far too slow to train with. It is completely independent of the backward pass, though, which is exactly what an audit needs.
The nudge. For parameter , hold every other parameter where it is and evaluate the loss twice: once with nudged up by , once nudged down by . The numerical gradient is
The comparison. Measure how far the claimed gradient is from with the relative error
taking when and are both exactly . Parameter is flagged when .
Task: write gradient_check(loss_fn, params, claimed, eps, tol).
loss_fn takes a list of parameter values, in the same order as params, and returns the loss as a number.params holds the current parameter values, and claimed holds the gradient a backward pass reported for each one.eps is and tol is the flagging threshold. Both are positive.[numerical, flagged]. numerical is the list of , each rounded to 4 decimal places. flagged lists the indices of the flagged parameters, counting from , in increasing order. Flag using the unrounded .In the tests, loss_fn is a small network's loss on fixed data, written as a lambda. For example
is a single ReLU neuron with weight p[0] and bias p[1], seeing the input . Its output is multiplied by an output weight p[2] and scored by squared error against a target of .
Framework authors test their own backward passes this way. They run the check once, on a tiny network, and then switch it off. A layer with ten thousand weights would need twenty thousand loss evaluations to check a single step, which is exactly the cost that one backward pass exists to avoid.