An Inception block runs four paths over the same input and joins them at the end by concatenating along the depth axis.
The input is a feature block of size . With "same" padding throughout, every path returns a grid.
| path | what it does | output channels |
|---|---|---|
| A | convolution | |
| B | convolution | |
| C | convolution | |
| D | max pooling, stride , then straight to the join | ? |
Pooling has nothing to learn: it takes a maximum inside each channel separately, so path D returns a block with exactly the depth it was handed.
What is the depth of this block's output, and what does it mean for a network that stacks blocks like this one?
Select all that apply.