A team mapping field boundaries in aerial photos wants an output exactly the size of the input, one label per pixel. So their network never pools, and every layer uses stride 1 with "same" padding. To let each output pixel still see a wide area, some layers use dilated filters.
A filter with dilation keeps its usual number of weights but spreads its taps out, so that neighbouring taps sit pixels apart instead of 1. Dilation 1 is ordinary convolution. For example, a 3×3 filter with dilation 2 has its taps at offsets , and from its centre in each direction. It covers a 5×5 square while reading only 9 of the pixels in it.
The network is this stack, applied in order to the photo:
| layer | filter size | dilation |
|---|---|---|
| 1 | 3×3 | 1 |
| 2 | 3×3 | 3 |
| 3 | 5×5 | 2 |
| 4 | 3×3 | 6 |
| 5 | 1×1 | 1 |
How many pixels wide is the receptive field of one unit in layer 5's output, measured on the input photo?
Give a whole number.