Pooling is the one stage of a convolutional network with nothing to learn in it. A small window sweeps the grid, and every stop is replaced by a single summary number — either the largest value lying under the window, or the mean of the values lying under it.
Task: write pool2d(grid, size, stride, mode), returning the pooled grid with every value rounded to 4 decimal places.
grid is a rectangular list of rows of numbers, and size is the side of the square window.stride is how far the window moves between stops, applied the same way across and down.mode is "max" (keep the largest value in the window) or "avg" (keep the mean of the window).stride == size sets them edge to edge, and a smaller stride makes consecutive windows overlap — legal, and it must work.The pooling layer you will actually meet is
size = 2, stride = 2, which halves each side and so quarters the number of positions. Everything else here is that same sweep with the knobs turned — and it is the identical sweep a convolution performs, which is why the two shrink a grid by exactly the same formula.