A convolution has two knobs beyond the filter itself. Padding puts a ring of zeros around the image, so the window has somewhere to sit when it reaches the border. Stride decides how far the window jumps between stops. Both change the size of what comes out.
The plain sweep — window in the top-left corner, one position at a time, stopping wherever it still fits — is taken as read here. Every case below turns at least one of the two knobs, and the work is in the ring and the jump rather than in the multiply-and-add.
Task: write conv2d(image, kernel, stride, padding), returning the feature map with every value rounded to 4 decimal places.
image is a rectangular list of rows of numbers. It is not always square, so the height and the width have to be counted separately.padding is the number of rows and columns of zeros added to each of the four sides. It happens first, and everything afterwards treats the padded rectangle as the image. padding = 0 is a legal call, and means the sweep runs on the image exactly as given.stride positions at a time, across and down. It may never hang off the padded edge. A side of length therefore offers stops, and a final position that would overhang is dropped rather than nudged back to fit.kernel is square, and may be even-sized as well as odd.With a 3×3 filter,
padding = 1andstride = 1, the output comes back exactly the size of the input — the "same" padding that makes deep stacks possible at all. Turn either knob and that no longer holds: the border grows the map, the jump shrinks it, and every architecture you will read about is a choice about which of the two is winning.