如何删除DataFrame中与其他行存在包含关系的冗余行?
Pandas删除被包含的冗余行高效实现方法
首先统一处理数据集里的空值,把空字符串、NaN等格式统一替换为pandas的空值标记,避免不同空格式判断出错:
import pandas as pd import numpy as np df = df.replace(['', np.nan], pd.NA)
场景1:空值仅出现在固定少数列(如示例中的topic列)
这是绝大多数业务场景的情况,实现逻辑简单,性能极高,可处理千万级以下数据集:
# 示例数据构造 df = pd.DataFrame({ 'id': [1008068494, 1008068494, 1008068494], 'name': ['abc', 'abc', 'abc'], 'label': ['x', 'x', 'y'], 'topic': ['animal', pd.NA, pd.NA], 'date': [20210929, 20210929, 20210929] }) # 填入除可能为空的列之外的所有关联列 key_cols = ['id', 'name', 'label', 'date'] # 按关联列分组,每组优先保留空值更少的行 df = df.sort_values('topic', na_position='last').groupby(key_cols, as_index=False).first()
运行后即可得到预期结果:仅保留id、name、label、date相同且topic非空的行,删除topic为空的冗余行。
场景2:空值可能出现在任意列
如果空值没有固定列,可通过排序加特征匹配的方式实现,可处理百万级以下数据集:
# 新增辅助列统计每行空值数量,按空值数量升序排列,空值越少的行优先级越高 df['na_count'] = df.isna().sum(axis=1) df = df.sort_values('na_count').reset_index(drop=True) to_drop = [] seen = set() for idx, row in df.iterrows(): # 生成当前行所有非空列的键值对元组作为特征 row_feature = tuple((col, row[col]) for col in df.columns if pd.notna(row[col])) # 检查当前行是否被已保留的行包含 included = False for exist_feature in seen: if all(item in exist_feature for item in row_feature): included = True break if included: to_drop.append(idx) else: seen.add(row_feature) # 删除冗余行和辅助列 df = df.drop(to_drop).drop('na_count', axis=1).reset_index(drop=True)
如果数据量超过千万级,可以将上述匹配逻辑替换为MinHash或局部敏感哈希优化匹配速度,普通业务场景下上述两种实现完全够用。
内容的提问来源于stack exchange,提问作者marlon
相关产品推荐
相关产品推荐

