To build a training example for something that happened last Tuesday, you need the feature values as they were last Tuesday — not as they are now. Grab today's values and your model trains on information that didn't exist at prediction time, scores brilliantly in testing, and collapses in production. That's label leakage, and it is the single most common way an ML project quietly fails.
A feature store prevents it with point-in-time correct lookups.
Task: write point_in_time_lookup(feature_log, requests) returning one value per request.
feature_log is a list of [entity, timestamp, value] entries, in no particular order.requests is a list of [entity, as_of] pairs.as_of, and return its value.as_of counts — it was already known at that moment.as_of must be ignored, however tempting. That's the whole point.None for that request.The discipline this encodes is what makes training data honest: a model that was going to run on Tuesday's information gets trained on Tuesday's information, so the score you measure offline is the score you'll actually get.