Matrix factorisation gives every user a short vector of hidden traits and every item a vector in the same space, then predicts a rating as their dot product. Training is one tiny correction per observed rating.
For the known rating r of item by user:
Task: write sgd_step(P, Q, user, item, rating, lr, reg) returning [P, Q] — both full matrices after the update, every value rounded to 4 decimal places.
P is a list of user vectors, Q a list of item vectors, all the same length.P_u in the Q_i update. Both updates are defined at the same instant, so overwriting P_u first and then feeding the new value into the Q_i update computes something else entirely. Compute both from the old values, then assign.reg may be 0, which turns the penalty off.The two update rules are mirror images, and that symmetry is the whole method: each factor moves in the direction that would have reduced this one error, scaled by how much the other factor contributed to it. A trait the item has strongly gets the biggest correction in the user's vector, and vice versa.
The - λP_u term is the same weight decay as in ridge regression, and it's load-bearing here. With most of the matrix unobserved, nothing stops the factors growing large enough to fit the few ratings a user has given exactly, and the penalty is what keeps them small enough to generalise.