Before any collaborative filtering, you can predict a surprising amount from three numbers: how high ratings run in general, whether this user is generous, and whether this item is well liked.
where μ is the global mean rating, and the two biases are each a deviation from it:
b_u = (this user's mean rating) − μb_i = (this item's mean rating) − μTask: write baseline_predict(ratings, user, item) returning the prediction, rounded to 4 decimal places.
ratings is a list of rows, one per user, one column per item, with None for unrated.μ is the mean over every known rating in the matrix.Nones in both.The subtractions are what makes this additive. Writing it as (user mean) + (item mean) would count the global level twice and give predictions roughly double what they should be — expressing each part as a deviation is what lets you add them to one shared baseline.
This is the model that wins more of the Netflix Prize than people expect. A generous user rating a well-liked film genuinely does rate it high, and nothing about their particular taste was needed to know that. Which is why serious systems predict the baseline first and apply collaborative filtering to the residual — leaving the hard model to explain only what the easy one couldn't.