You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在同一DataFrame多行间应用模糊匹配实现近似去重?

Pandas DataFrame行间模糊匹配去重实现

针对常规drop_duplicates无法处理的近似重复行(如"Rahul"与"Rahul m"),可以借助模糊匹配算法实现行间去重,以下是具体实现方案:

准备工作

首先安装所需依赖库:

pip install fuzzywuzzy python-Levenshtein

示例DataFrame构建

先还原你的示例数据:

import pandas as pd

df = pd.DataFrame({
    'ID': [1, 3, 1, 5, 6, 1],
    'name': ['Rahul', 'sarthak', 'Rahul m', 'priyansh', 'Aman', 'Rahul'],
    'salary': [7000, 5000, 7000, 2500, 3500, 7000]
})

实现逻辑

通过两两计算文本相似度,结合业务字段(如ID、salary)先分组,再在组内进行模糊匹配去重,既减少计算量又保证逻辑精准:

from fuzzywuzzy import fuzz

def remove_fuzzy_duplicates(group, threshold=90):
    # 初始化保留的行索引列表
    keep_indices = []
    # 遍历组内每一行
    for idx in group.index:
        current_name = group.loc[idx, 'name']
        # 检查当前行是否与已保留行的相似度超过阈值
        is_duplicate = False
        for keep_idx in keep_indices:
            compare_name = group.loc[keep_idx, 'name']
            if fuzz.ratio(current_name, compare_name) >= threshold:
                is_duplicate = True
                break
        if not is_duplicate:
            keep_indices.append(idx)
    return group.loc[keep_indices]

# 按ID和salary分组,组内执行模糊去重
deduplicated_df = df.groupby(['ID', 'salary'], group_keys=False).apply(remove_fuzzy_duplicates)

结果说明

执行代码后,deduplicated_df的输出为:

ID      name  salary
0   1     Rahul    7000
2   3   sarthak    5000
4   5  priyansh    2500
5   6      Aman    3500

可以看到,"Rahul"、"Rahul m"以及重复的"Rahul"被合并保留唯一行,实现了近似去重。

可选优化

  • 调整threshold参数:数值越低匹配越宽松,可根据业务需求灵活设置
  • 改用fuzz.partial_ratio:适用于匹配部分文本的场景(如"Rahul"和"Rahul Sharma")
  • 大数据集适配:用rapidfuzz替代fuzzywuzzy,大幅提升计算速度

内容的提问来源于stack exchange,提问作者abhijeet purandare

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 16:58:58