You have a few extreme values and two tempting bad options: delete those rows (losing everything else they contain) or leave them in (letting them dominate every mean and every fitted coefficient). Winsorization takes a third route — keep the rows, but clamp the extremes to a percentile cutoff.
Task: write winsorize(values, lower_pct, upper_pct) returning the clamped values, each rounded to 4 decimal places.
lower_pct percentile and at the upper_pct percentile.For both percentiles use the linear-interpolation rule on the sorted values:
interpolating between neighbours when the position isn't a whole number.
lower_pct is below upper_pct, and both lie between 0 and 100.lower_pct = 0 and upper_pct = 100 the cutoffs are the actual minimum and maximum, so nothing is clamped.The distinction from trimming is the whole idea. Trimming removes the extreme rows; winsorizing keeps them and only flattens the one offending feature. So the row's other features, and its label, stay in your training set — you've suppressed one value's leverage without discarding a real observation.
What it costs you is the information that something extreme happened. After winsorizing, a transaction of £5,000,000 and one of £50,000 may both read as the 90th-percentile value, indistinguishable. If being extreme is itself the signal — as in fraud detection — clamp at your peril, and keep a flag for "this value was clamped".