You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Presidio拒绝列表中带重音字符未被识别,YAML已UTF-8编码

Presidio带重音字符的拒绝列表无法匹配问题解决

问题说明

使用Microsoft Presidio进行文本分析与匿名化时,通过all-config.yml配置的带重音字符的拒绝列表无法匹配文本中的对应内容,即便YAML文件已保存为UTF-8编码。例如配置中的hétérosexuel无法被识别,导致该单词未被匿名化。

示例代码

from presidio_analyzer import AnalyzerEngineProvider
from presidio_anonymizer import AnonymizerEngine

FR_TEXT = """Nom complet : Jean Dupont  
Préférence sexuelle : Jean s'identifie comme hétérosexuel"""

analyzer_conf_file = "path/to/all-config.yml"

provider = AnalyzerEngineProvider(analyzer_engine_conf_file=analyzer_conf_file)
analyzer = provider.create_engine()

analyzer_results = analyzer.analyze(text=FR_TEXT, language="fr")
anonymizer = AnonymizerEngine()
result = anonymizer.anonymize(text=FR_TEXT, analyzer_results=analyzer_results)

print(result.text)

YAML配置文件(all-config.yml)

supported_languages: 
- en
- fr
- nl

default_score_threshold: 0

nlp_configuration:
  nlp_engine_name: spacy
  models:
  -
    lang_code: en
    model_name: en_core_web_lg
  -
    lang_code: fr
    model_name: fr_core_news_lg
  -
    lang_code: nl
    model_name: nl_core_news_lg

recognizer_registry:
  global_regex_flags: 26

  recognizers: 
  - name: "SexualityFr"
    supported_language: "fr"
    supported_entity: "SEXUALITY"
    deny_list: [hétérosexuel]
    deny_list_score: 1

问题原因

  1. 正则标志配置错误:当前设置的global_regex_flags: 26未启用Unicode感知的匹配规则,导致带重音的Unicode字符无法被正确识别。
  2. 字符归一化不一致:Presidio默认的DenyListRecognizer对文本和拒绝列表项做小写转换,但未统一Unicode字符归一化格式(如NFC/NFD),导致带重音字符匹配失效。

解决方案

方案1:修正正则全局标志

将recognizer_registry下的global_regex_flags修改为34,对应Python正则模块的re.IGNORECASE | re.UNICODE标志,确保匹配时支持Unicode字符的大小写忽略:

recognizer_registry:
  global_regex_flags: 34

  recognizers: 
  - name: "SexualityFr"
    supported_language: "fr"
    supported_entity: "SEXUALITY"
    deny_list: [hétérosexuel]
    deny_list_score: 1

方案2:统一Unicode字符归一化

确保拒绝列表和输入文本的字符采用相同的归一化格式(如NFC):

  • 在代码中对输入文本做归一化处理:
import unicodedata

FR_TEXT = unicodedata.normalize('NFC', """Nom complet : Jean Dupont  
Préférence sexuelle : Jean s'identifie comme hétérosexuel""")
  • 同时确保YAML配置中的拒绝列表项使用相同归一化格式的字符。

方案3:自定义带重音感知的拒绝列表识别器

如果上述方法无效,可自定义识别器,手动处理字符归一化和匹配逻辑:

from presidio_analyzer import DenyListRecognizer, AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
import unicodedata

def normalize_text(text):
    # 归一化并转为小写,确保匹配一致性
    return unicodedata.normalize('NFC', text.lower())

class AccentAwareDenyListRecognizer(DenyListRecognizer):
    def __init__(self, deny_list, supported_entity, supported_language="en"):
        super().__init__(deny_list=deny_list, supported_entity=supported_entity, supported_language=supported_language)
        self.normalized_deny_list = [normalize_text(item) for item in deny_list]

    def analyze(self, text, entities, nlp_artifacts=None):
        normalized_input = normalize_text(text)
        results = []
        for item in self.normalized_deny_list:
            start_pos = normalized_input.find(item)
            while start_pos != -1:
                end_pos = start_pos + len(item)
                # 构建识别结果,映射回原文本位置(此处假设归一化后文本长度与原文本一致)
                results.append(
                    self.build_result(start=start_pos, end=end_pos, score=self.deny_list_score)
                )
                start_pos = normalized_input.find(item, end_pos)
        return results

# 初始化分析器并添加自定义识别器
FR_TEXT = """Nom complet : Jean Dupont  
Préférence sexuelle : Jean s'identifie comme hétérosexuel"""

analyzer = AnalyzerEngine()
analyzer.registry.add_recognizer(
    AccentAwareDenyListRecognizer(
        deny_list=["hétérosexuel"],
        supported_entity="SEXUALITY",
        supported_language="fr"
    )
)

analyzer_results = analyzer.analyze(text=FR_TEXT, language="fr")
anonymizer = AnonymizerEngine()
result = anonymizer.anonymize(text=FR_TEXT, analyzer_results=analyzer_results)

print(result.text)

内容的提问来源于stack exchange,提问作者Sergiy Polovynko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 11:23:24