You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清理爬取数据中的重复名称与重复文本片段?

爬取文本重复内容的识别与删除方案

一、处理末尾缺空格导致的拼接式名称重复

针对类似Bakhtiar Mohammed AbdullaBakhtiar Mohammed Abdulla这类完全重复的名称拼接,可通过以下两种方式处理:

1. 通用正则匹配法

利用正则反向引用匹配连续重复的字符串块,直接提取单份内容:

import re

def remove_concat_duplicates(text):
    # 匹配整段完全重复的拼接内容
    full_match = re.match(r'^(.+)\1$', text.strip())
    if full_match:
        return full_match.group(1)
    # 处理文本中间的拼接重复片段
    return re.sub(r'(.+?)\1+', r'\1', text)

# 测试示例
sample_name = "Bakhtiar Mohammed AbdullaBakhtiar Mohammed Abdulla"
print(remove_concat_duplicates(sample_name))
# 输出: Bakhtiar Mohammed Abdulla

2. 结合NER的精准处理

如果spaCy能识别出PERSON实体,可对识别结果做二次校验,修复重复拼接:

import spacy

nlp = spacy.load("en_core_web_sm")

def fix_duplicate_name(entity_text):
    text_len = len(entity_text)
    if text_len % 2 == 0:
        half_len = text_len // 2
        first_half = entity_text[:half_len]
        second_half = entity_text[half_len:]
        if first_half == second_half:
            return first_half
    return entity_text

# 处理文本中的人名实体
doc = nlp("The suspect is Bakhtiar Mohammed AbdullaBakhtiar Mohammed Abdulla.")
for ent in doc.ents:
    if ent.label_ == "PERSON":
        fixed_name = fix_duplicate_name(ent.text)
        print(f"原名称: {ent.text} → 修正后: {fixed_name}")

二、处理整段文本重复

1. 段落级去重

将文本按段落拆分后,通过集合记录已出现的段落,仅保留首次出现的内容:

def remove_duplicate_paragraphs(text):
    paragraphs = [p.strip() for p in text.split('\n') if p.strip()]
    seen = set()
    unique_paragraphs = []
    for para in paragraphs:
        if para not in seen:
            seen.add(para)
            unique_paragraphs.append(para)
    return '\n'.join(unique_paragraphs)

# 测试示例
sample_text = """慕尼黑火车站人流量较大。
慕尼黑火车站人流量较大。
警方已加强巡逻。"""
print(remove_duplicate_paragraphs(sample_text))

2. 段落内重复内容清理

针对单段内的重复内容,用正则匹配重复块并替换:

def remove_inline_duplicates(text):
    # 匹配重复的句子或短语(非贪婪模式避免过度匹配)
    return re.sub(r'(.+?)(?:\s*\1)+', r'\1', text)

# 测试示例
sample_inline = "慕尼黑火车站人流量较大。慕尼黑火车站人流量较大。警方已加强巡逻。"
print(remove_inline_duplicates(sample_inline))

三、解决spaCy NER在PyCharm与Jupyter的识别差异

差异核心原因是环境依赖版本不一致,解决步骤:

  1. 在两个环境中分别执行以下命令,确认spaCy和模型版本完全一致:
    pip show spacy
    python -m spacy validate
    
  2. 如果版本不同,在PyCharm和Jupyter中统一安装相同版本:
    pip install spacy==[指定版本号]
    python -m spacy download en_core_web_sm
    
  3. 清除Jupyter内核缓存,重启内核后重新加载模型。

内容的提问来源于stack exchange,提问作者Linda Brck

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 21:40:25