The Matrix Factorisation lesson fitted every person's and every item's taste scores with gradient descent. There is another way to fit them, with no learning rate at all.
Freeze one side. Suppose the recipes' taste scores are fixed. Then each customer's predicted ratings are dot products of their unknown scores with numbers you already know. That is a linear regression: the recipes they rated are the rows, the recipes' scores are the features, and their ratings are the targets. So each customer's best scores can be solved for directly. Then freeze the customers and solve for each recipe the same way. Repeating these two half-steps is called alternating least squares.
A meal-kit company tries it. ratings is the user-item matrix: one row per customer, one column per recipe, ratings from 1 to 5, and None for a blank cell. recipe_scores holds each recipe's starting scores on hidden tastes, one list of numbers per recipe.
To stop scores from blowing up for customers with very few ratings, each solve adds a penalty , just as ridge regression does. For one customer, using only the recipes they rated, their new score vector solves
where is rated recipe 's score vector, is the customer's rating of it, and is the identity matrix. The matrix on the left is : its entry in row , column is , plus when . So this is a system of linear equations in unknowns. A recipe's update is the same equation with the roles swapped: sum over the customers who rated it, using their score vectors and their ratings of it.
One sweep is:
Task: write als(ratings, recipe_scores, lam, sweeps) that runs sweeps sweeps and returns (customer_scores, recipe_scores): two lists of lists of plain Python floats, every number rounded to 4 decimal places.
With , every system is a single equation. Customer 0 rated recipe 0 (5 stars) and recipe 2 (1 star), so , giving .