如何基于指定列删除Pandas DataFrame中的重复行?
Hey there, great question! You were actually on the right track with df.drop_duplicates(subset=['id'])—the missing piece is the keep parameter that lets you choose to keep the first or last occurrence of each duplicate group. Let me break this down clearly for you.
核心用法:指定列 + 控制保留哪条记录
The drop_duplicates() method is exactly the right tool here, and the subset parameter is how you tell pandas to only check for duplicates based on your target column (like id). The keep parameter then decides which duplicate entry to keep:
1. 保留每个唯一值的第一次出现
If you want to keep the first row for each unique id and drop all later duplicates, use keep='first' (this is actually the default behavior, so you might have been getting this without explicitly setting it, but it's good to write it out for clarity):
# 保留每个id的第一行,删除后续重复行 df_cleaned = df.drop_duplicates(subset=['id'], keep='first')
2. 保留每个唯一值的最后一次出现
To keep the most recent (last) entry for each id instead, just switch to keep='last':
# 保留每个id的最后一行,删除前面的重复行 df_cleaned = df.drop_duplicates(subset=['id'], keep='last')
3. 删除所有重复行(无保留)
If you want to remove every row that has a duplicate id (keeping only rows with completely unique ids), use keep=False:
# 删除所有存在重复id的行,只保留id唯一的行 df_cleaned = df.drop_duplicates(subset=['id'], keep=False)
直接修改原DataFrame(无需赋值新变量)
If you don't want to create a new DataFrame and instead want to modify the original one in place, add the inplace=True parameter:
# 直接在原DataFrame上修改,不需要赋值给新变量 df.drop_duplicates(subset=['id'], keep='first', inplace=True)
实际示例看效果
Let's use a sample DataFrame to make this concrete:
import pandas as pd # 示例DataFrame df = pd.DataFrame({ 'id': [1, 2, 2, 3, 3, 3], 'value': ['a', 'b', 'c', 'd', 'e', 'f'] })
Original DataFrame:
id value 0 1 a 1 2 b 2 2 c 3 3 d 4 3 e 5 3 f
After keep='first':
id value 0 1 a 1 2 b 3 3 d
After keep='last':
id value 0 1 a 2 2 c 5 3 f
关于效率的小提示
Rest assured, drop_duplicates() is the most efficient way to do this in pandas—it's optimized to handle large DataFrames quickly, so you don't need to mess with manual grouping or loops which would be much slower for big datasets.
内容来源于stack exchange

