One-hot encoding a column with ten thousand distinct values gives you ten thousand columns. Target encoding replaces each category with one number instead: the average target value for rows in that category.
That works beautifully for common categories and disastrously for rare ones. A category appearing once gets encoded as that single row's target — which is the target leaking straight into the feature.
Smoothing fixes it by blending each category's mean toward the global mean, weighted by how much evidence there is:
Task: write target_encode(categories, targets, smoothing) returning the encoding for each row, rounded to 4 decimal places.
categories[i] is row i's category and targets[i] its target value. Rows sharing a category share an encoding.n_c is how many rows that category has and mean_c their average target. mean_global is the average over all rows.smoothing (the m above) is zero or positive. At m = 0 you get the raw category mean back, unsmoothed.Read the formula as a tug-of-war between two weights. A category with 500 rows and m = 10 barely moves — the evidence overwhelms the prior. A category with 1 row and m = 10 is pulled almost all the way to the global mean, which is the honest answer: you know nothing specific about it.
One thing this version doesn't do, and which production code must: computing the means from the same rows the encoding will be used on still leaks, even smoothed. Real pipelines compute the encoding out-of-fold — each row encoded using the other folds' statistics only — so a row never contributes to its own feature.