You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

搭配spaCy引擎的Presidio为何无法识别ORG与PL_PESEL?

问题:Presidio结合spaCy波兰语模型无法识别ORGANIZATION和PL_PESEL实体

我使用spaCy的pl_core_news_lg模型提取波兰语文本中的命名实体,可正确识别persName(人名)和orgName(机构名),代码及输出如下:

import spacy

nlp = spacy.load("pl_core_news_lg")
text = "Jan Kowalski pracuje w IBM i współpracuje z Microsoft oraz Google."

doc = nlp(text)
entities = [(ent.text, ent.label_) for ent in doc.ents]

print(entities)

输出:

[('Jan Kowalski', 'persName'), ('IBM', 'orgName'), ('Microsoft', 'orgName'), ('Google', 'orgName')]

但使用同模型的Presidio及配置文件时,尽管支持实体列表包含ORGANIZATION和PL_PESEL,却无法正确识别这两类实体,代码、输出及配置文件如下:

from presidio_analyzer import AnalyzerEngine, RecognizerRegistry
from presidio_analyzer.nlp_engine import NlpEngineProvider

provider = NlpEngineProvider(conf_file="path_to_my_file/nlp_config.yaml") 
nlp_engine = provider.create_engine()

print(f"Supported recognizers (from NLP engine): {nlp_engine.get_supported_entities()}")

supported_languages = list(nlp_engine.get_supported_languages())
registry = RecognizerRegistry(supported_languages=["pl"])
registry.load_predefined_recognizers(["pl"])

print(f"Supported recognizers (from registry): {registry.get_supported_entities(['pl'])}")

analyzer = AnalyzerEngine(
    registry=registry, supported_languages=supported_languages, nlp_engine=nlp_engine
)

results = analyzer.analyze(text, "pl")

for entity in results:
    print(f"Found entity: {entity.entity_type} with score {entity.score}")

输出:

Supported recognizers (from NLP engine): ['ID', 'NRP', 'DATE_TIME', 'PERSON', 'LOCATION']
Supported recognizers (from registry): ['IN_VOTER', 'URL', 'IBAN_CODE', 'CREDIT_CARD', 'DATE_TIME', 'NRP', 'PHONE_NUMBER', 'MEDICAL_LICENSE', 'PERSON', 'IP_ADDRESS', 'ORGANIZATION', 'CRYPTO', 'LOCATION', 'PL_PESEL', 'EMAIL_ADDRESS']

配置文件:

nlp_engine_name: spacy
models:
  - lang_code: pl
    model_name: pl_core_news_lg

ner_model_configuration:
  model_to_presidio_entity_mapping:
    persName: PERSON
    orgName: ORGANIZATION
#    orgName: ORG
    placeName: LOCATION
    geogName: LOCATION
    LOC: LOCATION
    GPE: LOCATION
    FAC: LOCATION
    DATE: DATE_TIME
    TIME: DATE_TIME
    NORP: NRP
    ID: ID

请问为何Presidio无法识别ORGANIZATION和PL_PESEL,而spaCy却可以?


问题原因分析

1. ORGANIZATION无法识别的原因

你的Presidio配置文件存在层级错误:ner_model_configuration 必须嵌套在对应语言的 model 条目下,而非与 models 平级。当前配置中,NLP引擎未读取到 orgName: ORGANIZATION 的映射规则,导致 nlp_engine.get_supported_entities() 的输出里不包含 ORGANIZATION,自然无法识别机构实体。

修正后的配置文件如下:

nlp_engine_name: spacy
models:
  - lang_code: pl
    model_name: pl_core_news_lg
    # 将实体映射配置嵌套到当前模型条目下
    ner_model_configuration:
      model_to_presidio_entity_mapping:
        persName: PERSON
        orgName: ORGANIZATION
        placeName: LOCATION
        geogName: LOCATION
        LOC: LOCATION
        GPE: LOCATION
        FAC: LOCATION
        DATE: DATE_TIME
        TIME: DATE_TIME
        NORP: NRP
        ID: ID

2. PL_PESEL无法识别的原因

你的测试文本完全不包含符合PL_PESEL格式的内容。PL_PESEL是波兰的11位数字身份识别码(例如99010100000),而你使用的测试文本仅包含人名和机构名,没有PESEL号码,因此Presidio无法识别该实体。


内容的提问来源于stack exchange,提问作者Maltion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 15:15:57