如何清理爬取数据中的重复名称与重复文本片段?
爬取文本重复内容的识别与删除方案
一、处理末尾缺空格导致的拼接式名称重复
针对类似Bakhtiar Mohammed AbdullaBakhtiar Mohammed Abdulla这类完全重复的名称拼接,可通过以下两种方式处理:
1. 通用正则匹配法
利用正则反向引用匹配连续重复的字符串块,直接提取单份内容:
import re def remove_concat_duplicates(text): # 匹配整段完全重复的拼接内容 full_match = re.match(r'^(.+)\1$', text.strip()) if full_match: return full_match.group(1) # 处理文本中间的拼接重复片段 return re.sub(r'(.+?)\1+', r'\1', text) # 测试示例 sample_name = "Bakhtiar Mohammed AbdullaBakhtiar Mohammed Abdulla" print(remove_concat_duplicates(sample_name)) # 输出: Bakhtiar Mohammed Abdulla
2. 结合NER的精准处理
如果spaCy能识别出PERSON实体,可对识别结果做二次校验,修复重复拼接:
import spacy nlp = spacy.load("en_core_web_sm") def fix_duplicate_name(entity_text): text_len = len(entity_text) if text_len % 2 == 0: half_len = text_len // 2 first_half = entity_text[:half_len] second_half = entity_text[half_len:] if first_half == second_half: return first_half return entity_text # 处理文本中的人名实体 doc = nlp("The suspect is Bakhtiar Mohammed AbdullaBakhtiar Mohammed Abdulla.") for ent in doc.ents: if ent.label_ == "PERSON": fixed_name = fix_duplicate_name(ent.text) print(f"原名称: {ent.text} → 修正后: {fixed_name}")
二、处理整段文本重复
1. 段落级去重
将文本按段落拆分后,通过集合记录已出现的段落,仅保留首次出现的内容:
def remove_duplicate_paragraphs(text): paragraphs = [p.strip() for p in text.split('\n') if p.strip()] seen = set() unique_paragraphs = [] for para in paragraphs: if para not in seen: seen.add(para) unique_paragraphs.append(para) return '\n'.join(unique_paragraphs) # 测试示例 sample_text = """慕尼黑火车站人流量较大。 慕尼黑火车站人流量较大。 警方已加强巡逻。""" print(remove_duplicate_paragraphs(sample_text))
2. 段落内重复内容清理
针对单段内的重复内容,用正则匹配重复块并替换:
def remove_inline_duplicates(text): # 匹配重复的句子或短语(非贪婪模式避免过度匹配) return re.sub(r'(.+?)(?:\s*\1)+', r'\1', text) # 测试示例 sample_inline = "慕尼黑火车站人流量较大。慕尼黑火车站人流量较大。警方已加强巡逻。" print(remove_inline_duplicates(sample_inline))
三、解决spaCy NER在PyCharm与Jupyter的识别差异
差异核心原因是环境依赖版本不一致,解决步骤:
- 在两个环境中分别执行以下命令,确认spaCy和模型版本完全一致:
pip show spacy python -m spacy validate - 如果版本不同,在PyCharm和Jupyter中统一安装相同版本:
pip install spacy==[指定版本号] python -m spacy download en_core_web_sm - 清除Jupyter内核缓存,重启内核后重新加载模型。
内容的提问来源于stack exchange,提问作者Linda Brck
相关产品推荐
相关产品推荐

