A factory sorts machine parts into pass () and fail () from a few measurements, using a linear support vector machine. For a row the model computes a score
and predicts when , and otherwise. The dividing line is , and the street around it runs from to .
Where a row sits. Its margin is , its score multiplied by its true label:
What training lowers. For one row, the cost is
The first part, the hinge loss, is zero for a row outside the street on its own side, and grows the deeper a row sits inside the street or past the line. The second part grows with the size of the weights. Small weights mean a wide street (its width is ), so this part keeps pushing towards the widest street. The number sets how hard it pushes.
One step of stochastic gradient descent, for one row. Work out the row's margin using the current and . Then:
Every step shrinks the weights a little. Only a row inside the street or on the wrong side also pulls the line towards classifying it correctly. The bias is never shrunk.
Task: write train_svm(X, y, lr, lam, epochs, X_new).
X is a list of rows of numbers, and y gives each row's label, 1 or -1. lr is the learning rate and lam is .0.0. One epoch is one pass over the rows in the order given, taking one step per row, and each step uses the weights exactly as the previous step left them. Run epochs epochs; epochs may be 0.X_new with the final weights and bias.Return a tuple (w, b, predictions): w is the list of weights and b the bias, both rounded to 4 decimal places, and predictions is a list of 1 or -1, one per row of X_new. Predict with the unrounded values.