You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升Spacy命名实体识别效果?it_core_news_lg输出优化求助

修复spaCy意大利语NER模型(it_core_news_lg)实体识别错误的方案

针对你遇到的实体附带多余字符(如?、Sì)、复合实体拆分错误等问题,可通过以下几种方式解决:

  • 预处理输入文本,清理干扰字符
    错误实体大多因原始文本中的标点、感叹词干扰导致,先做文本清洗:

    • 用正则移除实体前后的非字母干扰字符:
      import re
      # 移除附着在单词前后的问号
      cleaned_text = re.sub(r'([a-zA-ZÀ-ÿ]+)\?+', r'\1', raw_text)
      cleaned_text = re.sub(r'\?+([a-zA-ZÀ-ÿ]+)', r'\1', cleaned_text)
      # 移除无关感叹词(如Sì)
      cleaned_text = re.sub(r'\bSì\b', '', cleaned_text)
      
  • 实体后处理过滤与拆分
    识别完成后对结果做二次修正:

    filtered_entities = set()
    nlp = spacy.load("it_core_news_lg")
    doc = nlp(cleaned_text)
    
    for ent in doc.ents:
        ent_text = ent.text.strip()
        # 拆分含连接词"e"的复合实体
        if ' e ' in ent_text:
            filtered_entities.update([name.strip() for name in ent_text.split(' e ')])
        else:
            filtered_entities.add(ent_text)
    
  • 微调模型适配特定场景
    如果通用模型无法适配你的文本风格,用自定义标注数据微调:

    import spacy
    from spacy.training import Example
    
    nlp = spacy.load("it_core_news_lg")
    optimizer = nlp.resume_training()
    
    # 标注包含错误案例的训练数据
    TRAIN_DATA = [
        ("Carlo Rossi?", {"entities": [(0, 11, "PER")]}),
        ("Bruno?Sì", {"entities": [(0, 5, "PER")]}),
        ("Maria e Pia", {"entities": [(0, 5, "PER"), (8, 11, "PER")]}),
        # 补充更多场景的标注数据
    ]
    
    for epoch in range(10):
        losses = {}
        for text, annotations in TRAIN_DATA:
            doc = nlp.make_doc(text)
            example = Example.from_dict(doc, annotations)
            nlp.update([example], sgd=optimizer, losses=losses)
        print(f"Epoch {epoch+1} Loss: {losses}")
    
  • 规则匹配增强NER
    用spaCy的PhraseMatcher添加精准人名规则,覆盖模型识别盲区:

    from spacy.matcher import PhraseMatcher
    
    nlp = spacy.load("it_core_news_lg")
    matcher = PhraseMatcher(nlp.vocab)
    # 加入你需要精准匹配的人名短语
    person_patterns = [nlp(name) for name in ["Carlo Rossi", "Pietro", "Bruno", "Maria", "Pia"]]
    matcher.add("TARGET_PERSONS", person_patterns)
    
    doc = nlp(raw_text)
    # 将匹配到的短语标记为PER实体
    for match_id, start, end in matcher(doc):
        span = doc[start:end]
        span.label_ = "PER"
    

内容的提问来源于stack exchange,提问作者Bambargiya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 06:30:22