A toy regression network, small enough to train by hand, takes two inputs, has two ReLU hidden neurons, and has one output neuron with no activation:
h1=ReLU(w11x1+w12x2+b1),h2=ReLU(w21x1+w22x2+b2)
y^=v1h1+v2h2+c
Its nine parameters currently are:
| neuron | first incoming weight | second incoming weight | bias |
|---|
| hidden 1 (fed by x1,x2) | w11=0.5 | w12=0.25 | b1=0 |
| hidden 2 (fed by x1,x2) | w21=−0.5 | w22=0.75 | b2=0.5 |
| output (fed by h1,h2) | v1=1 | v2=−0.5 | c=0.25 |
The training example is x1=1, x2=2 with target y=2. The loss is the squared error on this one example, L=(y^−y)2.
Run one complete training step on this example: forward, loss, backward, update. The update is plain gradient descent on every weight and bias,
p←p−η∂p∂L,η=0.02
Then run the forward pass again on the same example, using the updated parameters.
What is the loss now?
Round to 3 decimal places.