如何使用spaCy在Python中精准识别分词的实体类型?
如何使用spaCy在Python中精准识别分词的实体类型?
看起来你在使用spaCy处理支付账单类文本的实体识别时,遇到了几个和预期不符的识别结果,我来帮你一步步优化这些问题:
一、先从更换预训练模型入手(最快速的优化)
你当前使用的en_core_web_sm是spaCy的轻量小模型,它的训练数据量有限,对专业领域文本(比如支付账单)的实体识别精度会打折扣。建议换成中等规模的en_core_web_md或者大规模的en_core_web_lg,这两个模型在实体识别的召回率和准确率上都会好很多。
修改后的代码示例:
import spacy # 换成中等规模模型,需要先安装:pip install en_core_web_md nlp = spacy.load("en_core_web_md") def getPayeeName(description): description = description.replace("-", " ").replace("/", " ").strip() doc = nlp(description) for token in doc: print(f"Token: {token.text}, Entity: {token.ent_type_ if token.ent_type_ else 'None'}") # Example input description = "UPI DR 400874707203 BENGALORE 08 JAN 2024 14:38:56 MEDICAL LTD HDFC 50200" getPayeeName(description)
更换模型后,大概率能修复BENGALORE未被识别为GPE的问题,同时对50200这类数字的误识别也会减少。
二、用自定义规则修正领域特有问题(针对支付账单场景)
如果更换模型后还有残留问题,比如UPI、DR被误标为ORG,或者某些数字串的识别错误,可以用spaCy的EntityRuler来添加自定义规则,强制修正这些实体:
带自定义规则的代码示例:
import spacy from spacy.pipeline import EntityRuler nlp = spacy.load("en_core_web_md") # 添加自定义实体规则,设置overwrite_ents=True让规则优先覆盖原NER结果 ruler = EntityRuler(nlp, overwrite_ents=True) # 规则1:标记BENGALORE为GPE ruler.add_patterns([{"label": "GPE", "pattern": "BENGALORE"}]) # 规则2:将UPI标记为自定义的支付系统类型,避免被识别为ORG ruler.add_patterns([{"label": "PAYMENT_SYSTEM", "pattern": "UPI"}]) # 规则3:将DR标记为交易类型,避免被识别为ORG ruler.add_patterns([{"label": "TRANSACTION_TYPE", "pattern": "DR"}]) # 规则4:长度为5的纯数字串标记为None,覆盖ORG误标 ruler.add_patterns([{"label": "None", "pattern": [{"SHAPE": "ddddd"}]}]) # 将规则器添加到pipeline中,放在NER之前生效 nlp.add_pipe(ruler, before="ner") def getPayeeName(description): description = description.replace("-", " ").replace("/", " ").strip() doc = nlp(description) for token in doc: ent_type = token.ent_type_ if token.ent_type_ != "None" else "None" print(f"Token: {token.text}, Entity: {ent_type}") # Example input description = "UPI DR 400874707203 BENGALORE 08 JAN 2024 14:38:56 MEDICAL LTD HDFC 50200" getPayeeName(description)
三、后处理手动修正(特殊场景的兜底方案)
如果某些极端情况还是无法通过模型或规则解决,你可以在遍历实体的时候,手动添加后处理逻辑,针对性修正:
- 检查如果实体类型是
ORG但token是纯数字,强制改为None - 检查特定的城市名、缩写,手动修正实体类型
后处理代码示例:
def getPayeeName(description): description = description.replace("-", " ").replace("/", " ").strip() doc = nlp(description) for token in doc: ent_type = token.ent_type_ if token.ent_type_ else "None" # 后处理1:纯数字的ORG改为None if ent_type == "ORG" and token.text.isdigit(): ent_type = "None" # 后处理2:BENGALORE手动改为GPE if token.text == "BENGALORE": ent_type = "GPE" # 后处理3:UPI、DR手动修正为非ORG类型 if token.text in ["UPI", "DR"]: ent_type = "None" # 也可以替换为自定义类型,比如PAYMENT_SYSTEM print(f"Token: {token.text}, Entity: {ent_type}")
四、长期最优方案:领域自定义NER训练
如果你的业务场景大量处理这类支付账单文本,建议收集一批标注好的支付账单数据,用spaCy训练一个自定义的NER模型。这样模型会完全适配你的领域文本,识别精度能达到最优,不过这个方案需要一定的数据标注和训练成本。
备注:内容来源于stack exchange,提问作者PrakashT
相关产品推荐
相关产品推荐

