The filter values are weights, and training finds them by gradient descent. This is where the scale of the pixels starts to matter.
Take one filter scoring one patch, with no bias:
where are the nine pixel values under the filter. The filter is trained on this one patch alone, toward a fixed target , with the loss
It uses plain gradient descent, with all nine weights updated together at every step:
The patch, as it comes from the camera:
| col 0 | col 1 | col 2 | |
|---|---|---|---|
| row 0 | 204 | 204 | 204 |
| row 1 | 153 | 153 | 153 |
| row 2 | 51 | 51 | 51 |
The filter is trained twice, both times from the same starting weights (which leave the score away from ) and toward the same target:
Each run has a critical learning rate. Below it, the gap between score and target shrinks at every step. Above it, the gap grows at every step and training blows up.
How many times larger is Run B's critical learning rate than Run A's?
Give the exact value. It is a whole number.