You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Dataprep中删除重复行及其首次出现的记录?

How to Remove Duplicate Rows (Including Their First Occurrence) in Google Dataprep

Got it, let's solve this—you need to eliminate every instance of rows that appear more than once, right? That means if a row duplicates anywhere in your dataset, both the first occurrence and all copies get deleted, leaving only rows that are 100% unique. Here's a step-by-step approach tailored for Google Dataprep:

  • Step 1: Calculate the occurrence count for each row group
    Add a new derived column to your dataset. Use this formula to count how many times each unique row (defined by your target columns) appears:

    COUNT(*) OVER (PARTITION BY col1, col2)
    

    Name this column something like row_occurrences. For your example data, this will assign:

    • 2 to both john,simpson rows
    • 1 to will,farrell and elon,musk
  • Step 2: Filter out rows with multiple occurrences
    Add a filter step to your recipe. Set the condition to keep only rows where:

    row_occurrences == 1
    

    This will automatically drop every row that has a duplicate (including the first one), leaving only your fully unique records.

Example Verification

Original data:
col1,col2
john,simpson
will,farrell
john,simpson
elon,musk

After applying the steps:
col1,col2
will,farrell
elon,musk

Pro Tip

If your "duplicate" definition includes more columns (not just col1 and col2), simply add all relevant columns to the PARTITION BY clause. For example, if you care about col3 too, the formula becomes:

COUNT(*) OVER (PARTITION BY col1, col2, col3)

内容的提问来源于stack exchange,提问作者Daniel Hennig

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:41:53