You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何删除DataFrame中与其他行存在包含关系的冗余行?

Pandas删除被包含的冗余行高效实现方法

首先统一处理数据集里的空值,把空字符串、NaN等格式统一替换为pandas的空值标记,避免不同空格式判断出错:

import pandas as pd
import numpy as np

df = df.replace(['', np.nan], pd.NA)

场景1:空值仅出现在固定少数列(如示例中的topic列)

这是绝大多数业务场景的情况,实现逻辑简单,性能极高,可处理千万级以下数据集:

# 示例数据构造
df = pd.DataFrame({
    'id': [1008068494, 1008068494, 1008068494],
    'name': ['abc', 'abc', 'abc'],
    'label': ['x', 'x', 'y'],
    'topic': ['animal', pd.NA, pd.NA],
    'date': [20210929, 20210929, 20210929]
})

# 填入除可能为空的列之外的所有关联列
key_cols = ['id', 'name', 'label', 'date']
# 按关联列分组,每组优先保留空值更少的行
df = df.sort_values('topic', na_position='last').groupby(key_cols, as_index=False).first()

运行后即可得到预期结果:仅保留id、name、label、date相同且topic非空的行,删除topic为空的冗余行。

场景2:空值可能出现在任意列

如果空值没有固定列,可通过排序加特征匹配的方式实现,可处理百万级以下数据集:

# 新增辅助列统计每行空值数量,按空值数量升序排列,空值越少的行优先级越高
df['na_count'] = df.isna().sum(axis=1)
df = df.sort_values('na_count').reset_index(drop=True)

to_drop = []
seen = set()
for idx, row in df.iterrows():
    # 生成当前行所有非空列的键值对元组作为特征
    row_feature = tuple((col, row[col]) for col in df.columns if pd.notna(row[col]))
    # 检查当前行是否被已保留的行包含
    included = False
    for exist_feature in seen:
        if all(item in exist_feature for item in row_feature):
            included = True
            break
    if included:
        to_drop.append(idx)
    else:
        seen.add(row_feature)

# 删除冗余行和辅助列
df = df.drop(to_drop).drop('na_count', axis=1).reset_index(drop=True)

如果数据量超过千万级,可以将上述匹配逻辑替换为MinHash或局部敏感哈希优化匹配速度,普通业务场景下上述两种实现完全够用。

内容的提问来源于stack exchange,提问作者marlon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 13:54:04