A convolution looks like a mistake. Its window covers a single pixel, so it sees no neighbours at all and cannot possibly detect an edge, a corner, or anything else spatial.
What it does see is the full depth — every feature map at that one position. It reads a column of numbers straight down through the block and returns a single number, and it does that at every position independently. The grid comes out exactly the size it went in; only the depth changes.
That is the operation Inception puts in front of its expensive paths, and the one a bottleneck block uses to squeeze 256 maps down to 64 before doing the work.
Task: write one_by_one(block, filters, biases), returning the output block as plain nested lists, rounded to 4 decimal places.
block is a list of feature maps: block[c][i][j] is the value of channel c at row i, column j. Every map has the same height and width.filters is a list of filters. Each filter is a list holding one weight per input channel — a filter spanning the full input depth.biases holds one bias per filter.Count what this costs: one weight per input channel per filter, and not one of them spatial. A path on a 256-deep block is over 800,000 weights; hand it a block this operation has already squeezed and that figure collapses by an order of magnitude, which is the only reason an Inception block is affordable at all.