A small bank wants a model that says "yes" or "no" to loan requests. Its table mixes several kinds of column, and each needs its own treatment.
| Column | Kind | What to do |
|---|---|---|
income | number, may be None | fill with the median, then scale |
debt | number, may be None | fill with the median, then scale |
job | word | one-hot, ignoring unseen jobs |
has_guarantor | already 0 or 1 | pass through untouched |
Task: write loan_desk(table, labels, new_rows, seed).
table and new_rows are dicts with those four columns. labels holds "yes" or "no" for each row of table.income and debt float so each None becomes np.nan.train_test_split(X, y, test_size=0.25, random_state=seed, stratify=y).ColumnTransformer with ("numbers", ..., ["income", "debt"]) and ("words", ..., ["job"]), using remainder="passthrough" so has_guarantor survives.Pipeline with KNeighborsClassifier(n_neighbors=3) last, and fit it on the training rows only.{"width": number of columns the ColumnTransformer outputs, "score": test score rounded to 3 decimal places, "predictions": list of predictions for new_rows}.