A tiny model has two weights and the loss
The lowest possible loss is , reached at and at .
Because of a bug in the initialisation code, starts at exactly . starts at . The model is trained with plain full-batch gradient descent at a learning rate of , with no momentum. The log:
| step | 0 | 5 | 10 | 20 | 30 | 50 | 100 | 200 |
|---|---|---|---|---|---|---|---|---|
| loss | 10.0000 | 4.1381 | 2.0942 | 1.1330 | 1.0162 | 1.0002 | 1.0000 | 1.0000 |
By step 200, both components of the gradient read .
The run is restarted from the same starting point with one change. Which change brings the loss down near ?
Select all that apply.