如何在同一DataFrame多行间应用模糊匹配实现近似去重?
Pandas DataFrame行间模糊匹配去重实现
针对常规drop_duplicates无法处理的近似重复行(如"Rahul"与"Rahul m"),可以借助模糊匹配算法实现行间去重,以下是具体实现方案:
准备工作
首先安装所需依赖库:
pip install fuzzywuzzy python-Levenshtein
示例DataFrame构建
先还原你的示例数据:
import pandas as pd df = pd.DataFrame({ 'ID': [1, 3, 1, 5, 6, 1], 'name': ['Rahul', 'sarthak', 'Rahul m', 'priyansh', 'Aman', 'Rahul'], 'salary': [7000, 5000, 7000, 2500, 3500, 7000] })
实现逻辑
通过两两计算文本相似度,结合业务字段(如ID、salary)先分组,再在组内进行模糊匹配去重,既减少计算量又保证逻辑精准:
from fuzzywuzzy import fuzz def remove_fuzzy_duplicates(group, threshold=90): # 初始化保留的行索引列表 keep_indices = [] # 遍历组内每一行 for idx in group.index: current_name = group.loc[idx, 'name'] # 检查当前行是否与已保留行的相似度超过阈值 is_duplicate = False for keep_idx in keep_indices: compare_name = group.loc[keep_idx, 'name'] if fuzz.ratio(current_name, compare_name) >= threshold: is_duplicate = True break if not is_duplicate: keep_indices.append(idx) return group.loc[keep_indices] # 按ID和salary分组,组内执行模糊去重 deduplicated_df = df.groupby(['ID', 'salary'], group_keys=False).apply(remove_fuzzy_duplicates)
结果说明
执行代码后,deduplicated_df的输出为:
ID name salary 0 1 Rahul 7000 2 3 sarthak 5000 4 5 priyansh 2500 5 6 Aman 3500
可以看到,"Rahul"、"Rahul m"以及重复的"Rahul"被合并保留唯一行,实现了近似去重。
可选优化
- 调整
threshold参数:数值越低匹配越宽松,可根据业务需求灵活设置 - 改用
fuzz.partial_ratio:适用于匹配部分文本的场景(如"Rahul"和"Rahul Sharma") - 大数据集适配:用
rapidfuzz替代fuzzywuzzy,大幅提升计算速度
内容的提问来源于stack exchange,提问作者abhijeet purandare
相关产品推荐
相关产品推荐

