Plain linear regression chases the training data as closely as it can. Give it correlated or noisy features and it responds with enormous coefficients that cancel each other out — a fit that is perfect on the training set and wild everywhere else.
Ridge regression adds a cost for being large. The closed form barely changes: add alpha along the diagonal before inverting.
Task: write fit_ridge(X, y, alpha) returning [intercept, coef_1, coef_2, ...], each rounded to 4 decimal places.
X with a 1.0 so the model has an intercept.alpha on every diagonal entry except the first one, which stays 0. Shrinking the intercept would pull your predictions toward zero instead of toward the data's own average, which is never what you want.alpha is zero or positive. At alpha = 0 you should get exactly the ordinary least-squares answer back.np.eye(n) builds the identity matrix.Turn alpha up and watch the coefficients shrink toward zero while the intercept drifts up to absorb what they no longer explain. That trade — a little bias bought with a lot less variance — is the whole idea.