A phone app's network has a convolutional layer that takes a 28 × 28 × 96 block and produces a 28 × 28 × 192 block, using 3 × 3 filters with "same" padding and stride 1.
- Standard version: 192 filters, each 3 × 3 and spanning the full input depth.
- Depthwise separable version, a trick used by many networks built for phones, splits the job into two steps:
- Depthwise: 96 filters of size 3 × 3, one per input channel. Each one slides over only its own channel, not the full depth, so this step outputs a 28 × 28 × 96 block.
- Pointwise: 192 filters of size 1 × 1, each spanning all 96 channels, which mix the channels into a 28 × 28 × 192 block.
At every position it visits, a filter does one multiplication per weight. Ignore biases.
How many fewer multiplications does the depthwise separable version need than the standard version, for one image? Give your answer in millions, rounded to 1 decimal place.