A ratings matrix is mostly holes — a typical user has rated a tiny fraction of the catalogue. Many algorithms (plain matrix multiplication, PCA, a dense neural layer) can't run on holes at all, so the simplest first move is to fill each one with that item's average rating.
Task: write impute_ratings(matrix) returning the filled matrix, every filled value rounded to 4 decimal places.
matrix is a list of rows, one per user, one column per item. None means unrated.None with the mean of its column — the average rating that item received from the users who did rate it.Filling by item rather than by user is the deliberate choice here. An item's mean captures something real and shared — a universally loved film has a high average — while a user's mean says only how generous they are, which carries no information about the specific item you're guessing.
What to be wary of: these filled values are invented, and nothing downstream can tell them apart from real ratings. The matrix now looks dense and confident when it isn't, so a model trained on it will be most certain exactly where it knows least. That's why serious systems prefer methods that work on the observed entries only — matrix factorisation trained on the known ratings alone, rather than on a filled-in matrix.