Stacking more layers onto a deep plain network made it worse — worse on the training set, which rules out overfitting and leaves an optimisation failure. The fix was one addition: carry a block's input past its layers and add it back, so the block computes instead of .
Here you will run a stack of those blocks over a single row of a feature map, so the whole thing stays small enough to follow by hand.
Task: write residual_stack(signal, kernels), returning the final row as a plain list, rounded to 4 decimal places.
signal is a list of numbers — one row of a feature map.kernels is a list of blocks, one 3-tap filter each.
One of the tests hands a block a filter of all zeros. In a plain stack that layer would wipe the row out and everything after it would be working with nothing; here it is forced to pass its input straight through. That is the whole argument for skip connections — a layer with nothing useful to add costs nothing instead of doing damage, which is what stopped extra depth from being a risk.