Regularization adds a charge for large weights to the loss. Gradient descent then feels two pulls on every weight: the data pulls it toward whatever fits the training set, and the penalty pulls it toward zero. This problem runs both pulls for several steps and tracks where the weights end up.
To make the data's pull concrete, each weight has a target , the value the training data on its own would choose for it. The data loss is a bowl centred on the targets:
The penalty is one of these, with strength (strength):
penalty | penalty term |
|---|---|
"l2" | |
"l1" |
Each step updates every weight in two stages, both using learning rate (lr):
The L1 penalty needs one extra rule. Its slope is wherever the weight is positive and wherever it is negative, so its stage moves a fixed distance toward zero. If that distance would reach or cross zero, the weight stops at exactly instead of going past it. If is exactly , it stays through the penalty stage.
Task: write penalised_descent(weights, targets, lr, strength, penalty, steps). Starting from weights, run steps full steps and return the final weights as a list of floats, rounded to 4 decimal places. weights and targets have the same length (at least 1), lr and strength are positive, penalty is "l1" or "l2", and steps is a whole number of at least 1.
Worked through, penalised_descent([1.0, 0.2], [0.8, 0.1], 0.5, 0.2, "l1", 2). The L1 stage moves each weight toward zero:
| step | weight | data stage | penalty stage |
|---|---|---|---|
| 1 | |||
| 1 | |||
| 2 | |||
| 2 | , so it stops at |
The answer is [0.7, 0.0].
Why stop at zero? Without that rule, the L1 stage subtracts the same fixed amount every step. A weight smaller than that amount jumps across zero, gets pushed back, and jumps again. It flickers around zero forever without ever being zero. The stop rule is what makes L1 switch weights off exactly, and exact zeros are the whole reason to choose L1 over L2.