Picture the thinnest deep network there is: one neuron per layer, stacked layers deep. Starting from , each layer computes
and the network's output is .
The backward pass starts at the output with . Every layer it crosses multiplies the signal by that layer's rate, which has two parts, the activation's slope and the weight:
The three activations and their slopes:
activation | ||
|---|---|---|
"sigmoid" | ||
"tanh" | ||
"relu" | if , otherwise (including at exactly ) |
Task: write stack_gradients(activation, x, weights, biases). Layer uses weights[l-1] and biases[l-1], and every layer uses the same activation. Run the forward pass, then the backward pass, and return the list
That is numbers: 1.0 at the output first, and the gradient that reaches the input last. Round each one to 4 decimal places. Read left to right, the list shows how much of the signal is still alive after each layer it crosses on the way back.