A help desk is building an FAQ from the questions its users have typed. Each question has already been turned into an embedding vector, so questions that mean the same thing point in nearly the same direction — however they were worded, and whatever their vectors' lengths.
Before the FAQ goes live, near-duplicates have to go. The rule is greedy, working through the questions in the order they arrived:
0 is always kept.threshold, it's a near-duplicate and is dropped. Otherwise it's kept.Task: write deduplicate(vectors, threshold), returning the indexes of the kept questions, in increasing order.
vectors is a non-empty list of vectors, all with the same number of components, and none of them all zeros.threshold is a number between 0 and 1.Example.
Questions 1 and 3 point almost exactly the way question 0 does, and question 4 almost exactly the way question 2 does — even though none of them has the same length as the question it repeats.