A model is trained on one view of the data and served another. The feature pipelines are supposed to match — but one team patches a unit conversion, or a currency field changes meaning, and suddenly the model is being fed numbers unlike anything it trained on. Accuracy falls and nothing has errored. This is train-serving skew, and you catch it by comparing feature statistics on both sides.
Task: write skew_report(train_stats, serve_stats, threshold) returning the sorted list of feature names that look wrong.
Both arguments are dictionaries mapping a feature name to its mean. Flag a feature if any of these holds:
train_stats but is missing from serve_stats — the model needs a feature serving isn't producing.serve_stats but not in train_stats — serving is producing something the model never saw.threshold, measured relatively:abs on the bottom too, so negative-valued features behave the same way as positive ones.The relative comparison is what makes one threshold work across every feature at once. A mean shifting from 30 to 31 is a shrug; the same absolute shift on a feature whose mean is 0.5 would be a catastrophe. Dividing by the training mean puts age in years and income in dollars on the same footing.