搭配spaCy引擎的Presidio为何无法识别ORG与PL_PESEL?
问题:Presidio结合spaCy波兰语模型无法识别ORGANIZATION和PL_PESEL实体
我使用spaCy的pl_core_news_lg模型提取波兰语文本中的命名实体,可正确识别persName(人名)和orgName(机构名),代码及输出如下:
import spacy nlp = spacy.load("pl_core_news_lg") text = "Jan Kowalski pracuje w IBM i współpracuje z Microsoft oraz Google." doc = nlp(text) entities = [(ent.text, ent.label_) for ent in doc.ents] print(entities)
输出:
[('Jan Kowalski', 'persName'), ('IBM', 'orgName'), ('Microsoft', 'orgName'), ('Google', 'orgName')]
但使用同模型的Presidio及配置文件时,尽管支持实体列表包含ORGANIZATION和PL_PESEL,却无法正确识别这两类实体,代码、输出及配置文件如下:
from presidio_analyzer import AnalyzerEngine, RecognizerRegistry from presidio_analyzer.nlp_engine import NlpEngineProvider provider = NlpEngineProvider(conf_file="path_to_my_file/nlp_config.yaml") nlp_engine = provider.create_engine() print(f"Supported recognizers (from NLP engine): {nlp_engine.get_supported_entities()}") supported_languages = list(nlp_engine.get_supported_languages()) registry = RecognizerRegistry(supported_languages=["pl"]) registry.load_predefined_recognizers(["pl"]) print(f"Supported recognizers (from registry): {registry.get_supported_entities(['pl'])}") analyzer = AnalyzerEngine( registry=registry, supported_languages=supported_languages, nlp_engine=nlp_engine ) results = analyzer.analyze(text, "pl") for entity in results: print(f"Found entity: {entity.entity_type} with score {entity.score}")
输出:
Supported recognizers (from NLP engine): ['ID', 'NRP', 'DATE_TIME', 'PERSON', 'LOCATION'] Supported recognizers (from registry): ['IN_VOTER', 'URL', 'IBAN_CODE', 'CREDIT_CARD', 'DATE_TIME', 'NRP', 'PHONE_NUMBER', 'MEDICAL_LICENSE', 'PERSON', 'IP_ADDRESS', 'ORGANIZATION', 'CRYPTO', 'LOCATION', 'PL_PESEL', 'EMAIL_ADDRESS']
配置文件:
nlp_engine_name: spacy models: - lang_code: pl model_name: pl_core_news_lg ner_model_configuration: model_to_presidio_entity_mapping: persName: PERSON orgName: ORGANIZATION # orgName: ORG placeName: LOCATION geogName: LOCATION LOC: LOCATION GPE: LOCATION FAC: LOCATION DATE: DATE_TIME TIME: DATE_TIME NORP: NRP ID: ID
请问为何Presidio无法识别ORGANIZATION和PL_PESEL,而spaCy却可以?
问题原因分析
1. ORGANIZATION无法识别的原因
你的Presidio配置文件存在层级错误:ner_model_configuration 必须嵌套在对应语言的 model 条目下,而非与 models 平级。当前配置中,NLP引擎未读取到 orgName: ORGANIZATION 的映射规则,导致 nlp_engine.get_supported_entities() 的输出里不包含 ORGANIZATION,自然无法识别机构实体。
修正后的配置文件如下:
nlp_engine_name: spacy models: - lang_code: pl model_name: pl_core_news_lg # 将实体映射配置嵌套到当前模型条目下 ner_model_configuration: model_to_presidio_entity_mapping: persName: PERSON orgName: ORGANIZATION placeName: LOCATION geogName: LOCATION LOC: LOCATION GPE: LOCATION FAC: LOCATION DATE: DATE_TIME TIME: DATE_TIME NORP: NRP ID: ID
2. PL_PESEL无法识别的原因
你的测试文本完全不包含符合PL_PESEL格式的内容。PL_PESEL是波兰的11位数字身份识别码(例如99010100000),而你使用的测试文本仅包含人名和机构名,没有PESEL号码,因此Presidio无法识别该实体。
内容的提问来源于stack exchange,提问作者Maltion
相关产品推荐
相关产品推荐

