如何在Pandas DataFrame中识别并输出近似重复的行?
检测科研文献标题的近似重复项(相似度≥95%)
核心前提:先做文本预处理
直接计算原标题的相似度容易被空格、大小写、标点干扰,这大概率是你之前用difflib没得到预期结果的原因。先统一清洗标题:
import pandas as pd from difflib import SequenceMatcher def clean_title(title): # 统一小写 title = title.lower() # 去除多余空格(包括换行、制表符) title = ' '.join(title.split()) # 移除无关标点(可根据学科需求调整保留项) punctuation = '!"#$%&\'()*+,-./:;<=>?@[\\]^_`{|}~' title = title.translate(str.maketrans('', '', punctuation)) return title # 为DataFrame添加清洗后的标题列 df['cleaned_title'] = df['Article'].apply(clean_title)
方案1:小数据集两两比对(精准)
如果你的文献数量在几千条以内,直接遍历所有两两组合,用SequenceMatcher计算相似度:
from itertools import combinations def get_near_duplicates(df, threshold=0.95): near_dup_list = [] # 遍历所有不重复的索引对 for idx1, idx2 in combinations(df.index, 2): t1 = df.loc[idx1, 'cleaned_title'] t2 = df.loc[idx2, 'cleaned_title'] # 计算相似度 sim_score = SequenceMatcher(None, t1, t2).ratio() if sim_score >= threshold: near_dup_list.append({ '原标题1': df.loc[idx1, 'Article'], '来源库1': df.loc[idx1, 'Database'], '原标题2': df.loc[idx2, 'Article'], '来源库2': df.loc[idx2, 'Database'], '相似度': round(sim_score, 4) }) return pd.DataFrame(near_dup_list) # 调用函数,阈值设为95% result_df = get_near_duplicates(df, threshold=0.95) print(result_df)
方案2:大数据集快速匹配(高效)
如果文献数量过万,两两比对效率太低,用fuzzywuzzy库实现快速批量匹配(需要先安装pip install fuzzywuzzy python-Levenshtein):
from fuzzywuzzy import fuzz, process def fast_near_duplicates(df, threshold=95): near_dup_list = [] cleaned_titles = df['cleaned_title'].tolist() processed_indices = set() for idx, title in enumerate(cleaned_titles): if idx in processed_indices: continue # 找出所有相似度≥阈值的标题 matches = process.extractBests( title, cleaned_titles, scorer=fuzz.ratio, score_cutoff=threshold ) # 只保留未处理过的、索引大于当前的匹配项(避免重复配对) for match_title, score, match_idx in matches: if match_idx > idx and match_idx not in processed_indices: near_dup_list.append({ '原标题1': df.loc[idx, 'Article'], '来源库1': df.loc[idx, 'Database'], '原标题2': df.loc[match_idx, 'Article'], '来源库2': df.loc[match_idx, 'Database'], '相似度': round(score/100, 4) }) processed_indices.add(match_idx) processed_indices.add(idx) return pd.DataFrame(near_dup_list) # 调用函数 result_df = fast_near_duplicates(df, threshold=95)
关键注意事项
- 预处理可灵活调整:比如医学文献中可能需要保留连字符、希腊字母,可修改标点移除规则;如果拼写错误较多,可额外添加拼写纠正(如
pyspellchecker库,但需注意专业术语的误修正) - 阈值可根据需求微调:95%的阈值适合仅存在微小差异的标题,若需覆盖更多拼写错误,可降到90%
- 若需合并重复项:可在找到近似对后,选择保留某一数据库的条目(如优先保留PubMed),或合并两个库的元数据
内容的提问来源于stack exchange,提问作者TrevorM
相关产品推荐
相关产品推荐

