You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中识别并输出近似重复的行?

检测科研文献标题的近似重复项(相似度≥95%)

核心前提:先做文本预处理

直接计算原标题的相似度容易被空格、大小写、标点干扰,这大概率是你之前用difflib没得到预期结果的原因。先统一清洗标题:

import pandas as pd
from difflib import SequenceMatcher

def clean_title(title):
    # 统一小写
    title = title.lower()
    # 去除多余空格(包括换行、制表符)
    title = ' '.join(title.split())
    # 移除无关标点(可根据学科需求调整保留项)
    punctuation = '!"#$%&\'()*+,-./:;<=>?@[\\]^_`{|}~'
    title = title.translate(str.maketrans('', '', punctuation))
    return title

# 为DataFrame添加清洗后的标题列
df['cleaned_title'] = df['Article'].apply(clean_title)

方案1:小数据集两两比对(精准)

如果你的文献数量在几千条以内,直接遍历所有两两组合,用SequenceMatcher计算相似度:

from itertools import combinations

def get_near_duplicates(df, threshold=0.95):
    near_dup_list = []
    # 遍历所有不重复的索引对
    for idx1, idx2 in combinations(df.index, 2):
        t1 = df.loc[idx1, 'cleaned_title']
        t2 = df.loc[idx2, 'cleaned_title']
        # 计算相似度
        sim_score = SequenceMatcher(None, t1, t2).ratio()
        if sim_score >= threshold:
            near_dup_list.append({
                '原标题1': df.loc[idx1, 'Article'],
                '来源库1': df.loc[idx1, 'Database'],
                '原标题2': df.loc[idx2, 'Article'],
                '来源库2': df.loc[idx2, 'Database'],
                '相似度': round(sim_score, 4)
            })
    return pd.DataFrame(near_dup_list)

# 调用函数,阈值设为95%
result_df = get_near_duplicates(df, threshold=0.95)
print(result_df)

方案2:大数据集快速匹配(高效)

如果文献数量过万,两两比对效率太低,用fuzzywuzzy库实现快速批量匹配(需要先安装pip install fuzzywuzzy python-Levenshtein):

from fuzzywuzzy import fuzz, process

def fast_near_duplicates(df, threshold=95):
    near_dup_list = []
    cleaned_titles = df['cleaned_title'].tolist()
    processed_indices = set()
    
    for idx, title in enumerate(cleaned_titles):
        if idx in processed_indices:
            continue
        # 找出所有相似度≥阈值的标题
        matches = process.extractBests(
            title, cleaned_titles, 
            scorer=fuzz.ratio, 
            score_cutoff=threshold
        )
        # 只保留未处理过的、索引大于当前的匹配项(避免重复配对)
        for match_title, score, match_idx in matches:
            if match_idx > idx and match_idx not in processed_indices:
                near_dup_list.append({
                    '原标题1': df.loc[idx, 'Article'],
                    '来源库1': df.loc[idx, 'Database'],
                    '原标题2': df.loc[match_idx, 'Article'],
                    '来源库2': df.loc[match_idx, 'Database'],
                    '相似度': round(score/100, 4)
                })
                processed_indices.add(match_idx)
        processed_indices.add(idx)
    return pd.DataFrame(near_dup_list)

# 调用函数
result_df = fast_near_duplicates(df, threshold=95)

关键注意事项

  • 预处理可灵活调整:比如医学文献中可能需要保留连字符、希腊字母,可修改标点移除规则;如果拼写错误较多,可额外添加拼写纠正(如pyspellchecker库,但需注意专业术语的误修正)
  • 阈值可根据需求微调:95%的阈值适合仅存在微小差异的标题,若需覆盖更多拼写错误,可降到90%
  • 若需合并重复项:可在找到近似对后,选择保留某一数据库的条目(如优先保留PubMed),或合并两个库的元数据

内容的提问来源于stack exchange,提问作者TrevorM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 07:05:30