A dataset has some rows that exactly repeat earlier rows. Before deciding whether to drop them, check if they actually change the answer you care about.
Task: write duplicate_check(column_names, rows, x, y). rows is a list of rows, each a list of values in the order of column_names. x and y are two of the column names.
Return a dictionary with:
"rows": how many rows there are"duplicates": how many rows repeat an earlier row in every column (the first copy doesn't count)"corr_all": the correlation between columns x and y using every row"corr_clean": the same correlation after removing the repeated rows"changed": True if the two correlations differ once rounded, otherwise FalseRound both correlations to 2 decimal places. There are no missing values, and neither x nor y holds the same value in every row, before or after removing repeats.