A tiny network reads one number and predicts one number. It has two hidden ReLU neurons and an output neuron with no activation:
Its parameters are
A teammate has hand-written the backward pass, which returns a gradient for all seven parameters. It has exactly one bug. For each hidden neuron's bias, the code reuses the line it has for that neuron's weight: it multiplies the neuron's error signal by the input . Everything else in the code is correct.
The teammate doesn't know about the bug yet. They plan to test the code with a gradient check on one example. For each of the seven parameters in turn, they nudge it up by and down by , run the forward pass both times, and compare
with the code's gradient for that parameter. On these examples no weighted sum gets anywhere near zero during the nudges, so the check measures every true slope exactly. It disagrees with the code precisely where the code's number is wrong.
They have four candidate examples:
| example | target | |
|---|---|---|
| 1 | ||
| 2 | ||
| 3 | ||
| 4 |
On which of these examples would the check show at least one disagreement?
Select all that apply.