Pipelines deliver duplicates. A retry re-sends a batch, two upstream systems both report the same event, a backfill overlaps a live load. Left alone, those copies quietly reweight your training data — the duplicated rows count twice, and the model leans toward them.
Task: write deduplicate(rows, keys) returning the rows with duplicates removed.
rows is a list of dictionaries. keys lists the field names that together identify a record.keys. The other fields are irrelevant to the comparison, however different they are."Keep the first" is a real decision and worth being deliberate about. With rows arriving oldest-first it keeps the original and discards corrections; keep the last instead and you treat later arrivals as updates. Either can be right — what's never right is leaving it unspecified, because then your pipeline's output depends on batch ordering nobody controls.
Build each row's identity as a tuple of its key values and track the ones you've seen in a set. Tuples are hashable where dictionaries and lists are not, which is precisely why this is the standard shape for a composite key.