A unit deep in a network only ever saw a small window of the layer below it. But each unit inside that window saw its own window of the layer below it, and so on down. Follow the chain all the way to the image and the region of pixels able to influence that one unit is far larger than any single filter.
A network is given as a list of layers, ordered from the input towards the output. Each layer is a pair [kernel, stride] — a square window of side kernel, moved stride positions at a time. Pooling layers are written the same way, since they sweep the same way: a 2×2 max pool with stride 2 is [2, 2]. There is no padding anywhere, and the input is large enough that the region never runs off its edge.
Task: write receptive_field(layers), returning a list whose -th entry is the side length — measured in input pixels — of the region that can affect a single unit in the output of layer .
receptive_field([[3, 1]]) gives [3.0], not [3].While every stride is 1, each layer widens the region by kernel - 1 pixels. One 3×3 layer reaches 3 pixels, a second reaches 5, a third 7.
A stride breaks that rule, and this is the part worth working through. Neighbouring units of a layer are built from windows one stride apart in whatever they read, so a stride-2 layer leaves its own units standing twice as far apart in the image as the units it read were. That spacing is the running product of every stride passed so far — after two stride-2 layers it is 4 input pixels, not 2. And a window of kernel units laid on units pixels apart reaches across input pixels, not kernel - 1.
Worked through, receptive_field([[3, 2], [3, 1]]):
[3.0, 7.0].This is the number that decides what a layer is even capable of recognising: a unit whose region covers twelve pixels cannot be reporting on an object forty pixels wide, no matter how it was trained.