如何在Google Dataprep中删除重复行及其首次出现的记录?
Got it, let's solve this—you need to eliminate every instance of rows that appear more than once, right? That means if a row duplicates anywhere in your dataset, both the first occurrence and all copies get deleted, leaving only rows that are 100% unique. Here's a step-by-step approach tailored for Google Dataprep:
Step 1: Calculate the occurrence count for each row group
Add a new derived column to your dataset. Use this formula to count how many times each unique row (defined by your target columns) appears:COUNT(*) OVER (PARTITION BY col1, col2)Name this column something like
row_occurrences. For your example data, this will assign:2to bothjohn,simpsonrows1towill,farrellandelon,musk
Step 2: Filter out rows with multiple occurrences
Add a filter step to your recipe. Set the condition to keep only rows where:row_occurrences == 1This will automatically drop every row that has a duplicate (including the first one), leaving only your fully unique records.
Example Verification
Original data:
col1,col2
john,simpson
will,farrell
john,simpson
elon,musk
After applying the steps:
col1,col2
will,farrell
elon,musk
Pro Tip
If your "duplicate" definition includes more columns (not just col1 and col2), simply add all relevant columns to the PARTITION BY clause. For example, if you care about col3 too, the formula becomes:
COUNT(*) OVER (PARTITION BY col1, col2, col3)
内容的提问来源于stack exchange,提问作者Daniel Hennig

