A hand-written network reads two numbers about a house, and , and predicts its price. It has two hidden ReLU neurons and one output neuron with no activation:
| hidden neuron | weight from | weight from | bias |
|---|---|---|---|
| 1 | |||
| 2 |
That makes nine parameters: four hidden weights, two hidden biases, two output weights and the output bias. The loss on one house is .
Like most hand-written networks, it keeps a cache. Every forward pass saves its inputs, each hidden neuron's weighted sum and output, and the prediction, overwriting whatever the previous forward pass saved. Whenever the backward pass needs a value from the forward pass, it reads it from the cache.
A training script does three things, in this order:
Nothing crashes, and nine gradients come back.
Which of them are correct for house A?
Select all that apply.