Two Inception blocks are stacked, one feeding the other. Every convolution uses "same" padding, so every path in a block returns the same grid size, and the paths are joined by concatenating along the depth axis.
Block A receives a 192-deep input:
| path | layers | filters in the final conv |
|---|---|---|
| one conv | 64 | |
| conv with 96 filters, then conv | 128 | |
| conv with 16 filters, then conv | 32 | |
| pool | max-pool (stride 1), then conv | 32 |
Block B takes block A's output as its input, and is built the same way:
| path | layers | filters in the final conv |
|---|---|---|
| one conv | 128 | |
| conv with 128 filters, then conv | 192 | |
| conv with 32 filters, then conv | 96 | |
| pool | max-pool (stride 1), then conv | 64 |
How many weights does block B hold?
Count convolution weights only — ignore biases, and remember pooling has none.
Give your answer in thousands, rounded to 1 decimal place (so 12,300 weights would be entered as 12.3).