如何从Pandas DataFrame中移除指定数量的重复行
移除DataFrame分组中最后n条重复行
你需要对DataFrame按指定列分组,移除每组中最后n条重复行(当组内行数超过n时保留前组内行数-n条,否则保留全部),可以通过以下方法实现:
核心实现代码
import pandas as pd # 原始DataFrame df = pd.DataFrame({ 'brand': ['Yum Yum', 'Yum Yum', 'Indomie', 'Indomie', 'Indomie'], }) n = 2 # 要移除的每组最后n条重复行 # 处理逻辑 filtered_df = df.groupby('brand').apply(lambda x: x.iloc[:-n] if len(x) > n else x).reset_index(drop=True) print(filtered_df)
输出结果
brand 0 Yum Yum 1 Yum Yum 2 Indomie
代码说明
groupby('brand'):按指定列(这里是brand)分组,将相同取值的行归为一组apply(lambda x: ...):对每个分组执行自定义判断:- 如果分组的总行数
len(x)大于n,用x.iloc[:-n]截取前len(x)-n条数据(即移除最后n条) - 如果分组行数小于等于n,直接保留整个分组
- 如果分组的总行数
reset_index(drop=True):重置结果的索引,避免分组后产生多级索引
多列重复的扩展用法
如果需要基于多列判断重复,只需修改groupby的列参数即可:
df = pd.DataFrame({ 'brand': ['Yum Yum', 'Yum Yum', 'Indomie', 'Indomie', 'Indomie'], 'flavor': ['chicken', 'chicken', 'beef', 'beef', 'beef'] }) n = 2 filtered_df = df.groupby(['brand', 'flavor']).apply(lambda x: x.iloc[:-n] if len(x) > n else x).reset_index(drop=True)
内容的提问来源于stack exchange,提问作者Gooby
相关产品推荐
相关产品推荐

