You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用spaCy在Python中精准识别分词的实体类型?

如何使用spaCy在Python中精准识别分词的实体类型?

看起来你在使用spaCy处理支付账单类文本的实体识别时,遇到了几个和预期不符的识别结果,我来帮你一步步优化这些问题:

一、先从更换预训练模型入手(最快速的优化)

你当前使用的en_core_web_sm是spaCy的轻量小模型,它的训练数据量有限,对专业领域文本(比如支付账单)的实体识别精度会打折扣。建议换成中等规模的en_core_web_md或者大规模的en_core_web_lg,这两个模型在实体识别的召回率和准确率上都会好很多。

修改后的代码示例:

import spacy

# 换成中等规模模型,需要先安装:pip install en_core_web_md
nlp = spacy.load("en_core_web_md")

def getPayeeName(description):
    description = description.replace("-", " ").replace("/", " ").strip()
    doc = nlp(description)

    for token in doc:
        print(f"Token: {token.text}, Entity: {token.ent_type_ if token.ent_type_ else 'None'}")

# Example input
description = "UPI DR 400874707203 BENGALORE 08 JAN 2024 14:38:56 MEDICAL LTD HDFC 50200"
getPayeeName(description)

更换模型后,大概率能修复BENGALORE未被识别为GPE的问题,同时对50200这类数字的误识别也会减少。

二、用自定义规则修正领域特有问题(针对支付账单场景)

如果更换模型后还有残留问题,比如UPI、DR被误标为ORG,或者某些数字串的识别错误,可以用spaCy的EntityRuler来添加自定义规则,强制修正这些实体:

带自定义规则的代码示例:

import spacy
from spacy.pipeline import EntityRuler

nlp = spacy.load("en_core_web_md")

# 添加自定义实体规则,设置overwrite_ents=True让规则优先覆盖原NER结果
ruler = EntityRuler(nlp, overwrite_ents=True)
# 规则1:标记BENGALORE为GPE
ruler.add_patterns([{"label": "GPE", "pattern": "BENGALORE"}])
# 规则2:将UPI标记为自定义的支付系统类型,避免被识别为ORG
ruler.add_patterns([{"label": "PAYMENT_SYSTEM", "pattern": "UPI"}])
# 规则3:将DR标记为交易类型,避免被识别为ORG
ruler.add_patterns([{"label": "TRANSACTION_TYPE", "pattern": "DR"}])
# 规则4:长度为5的纯数字串标记为None,覆盖ORG误标
ruler.add_patterns([{"label": "None", "pattern": [{"SHAPE": "ddddd"}]}])

# 将规则器添加到pipeline中,放在NER之前生效
nlp.add_pipe(ruler, before="ner")

def getPayeeName(description):
    description = description.replace("-", " ").replace("/", " ").strip()
    doc = nlp(description)

    for token in doc:
        ent_type = token.ent_type_ if token.ent_type_ != "None" else "None"
        print(f"Token: {token.text}, Entity: {ent_type}")

# Example input
description = "UPI DR 400874707203 BENGALORE 08 JAN 2024 14:38:56 MEDICAL LTD HDFC 50200"
getPayeeName(description)

三、后处理手动修正(特殊场景的兜底方案)

如果某些极端情况还是无法通过模型或规则解决,你可以在遍历实体的时候,手动添加后处理逻辑,针对性修正:

  • 检查如果实体类型是ORG但token是纯数字,强制改为None
  • 检查特定的城市名、缩写,手动修正实体类型

后处理代码示例:

def getPayeeName(description):
    description = description.replace("-", " ").replace("/", " ").strip()
    doc = nlp(description)

    for token in doc:
        ent_type = token.ent_type_ if token.ent_type_ else "None"
        # 后处理1:纯数字的ORG改为None
        if ent_type == "ORG" and token.text.isdigit():
            ent_type = "None"
        # 后处理2:BENGALORE手动改为GPE
        if token.text == "BENGALORE":
            ent_type = "GPE"
        # 后处理3:UPI、DR手动修正为非ORG类型
        if token.text in ["UPI", "DR"]:
            ent_type = "None"  # 也可以替换为自定义类型,比如PAYMENT_SYSTEM
        print(f"Token: {token.text}, Entity: {ent_type}")

四、长期最优方案:领域自定义NER训练

如果你的业务场景大量处理这类支付账单文本,建议收集一批标注好的支付账单数据,用spaCy训练一个自定义的NER模型。这样模型会完全适配你的领域文本,识别精度能达到最优,不过这个方案需要一定的数据标注和训练成本。


备注:内容来源于stack exchange,提问作者PrakashT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 17:59:51