A gym wants to phone members who are about to quit and offer them a discount before they go. At midnight on the 1st of every month, a model will score each active member: will this member cancel during the coming month?
The training table has one row per member per month from last year, labelled with whether that member cancelled in that month. The first model scores 97% accuracy on the validation set, which is suspiciously good. So the team trains a separate one-feature model on each candidate feature to see where the accuracy comes from:
| feature | what it records | accuracy on its own |
|---|---|---|
days_since_last_visit | days between the member's last check-in and the 1st | 91% |
contract_months_left | months still to run on the member's contract, as of the 1st | 74% |
locker_key_returned | yes if the member handed their locker key back at the front desk during the month (members return it on their final visit) | 86% |
cancelled_before | yes if the member had cancelled and later rejoined at any point before this month | 68% |
Which feature is target leakage, and has to be removed before the model goes live?
Select all that apply.