You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python模糊匹配为XML标记persName标签的失效问题排查

模糊匹配无法生效的原因及修复方案

问题背景

我有两份XML文件:第一份aix_xml_raw中的人名已用<persName>标签完成标记;第二份wien_xml_raw内容相近,但存在拼写差异且新增了部分段落。我希望通过模糊匹配(例如第一份中的mr. l Conte de Sle匹配第二份的mr. le C. de Sli.),将第一份中<persName>元素的内容在第二份中定位并标记,但添加模糊匹配的判断条件后代码无法生效,移除该条件后反而能正常运行,请问这是为什么?

原始XML示例

aix_xml_raw = """
    <doc><p>Louërent Dieu, mais <persName>mr. l C<ex>onte</ex> de Sle.</persName> ne voulut pas
        appliquer mon voyage a mon avantage il crut que cela ressembloit fort a l’avanture,
        et que la peur de me confesser au <persName>RP. Br.</persName> m’avoit fait aller a 
        <placeName>Hilzing</placeName> 
        Je ne m’excuse point, laissant au jugement de ceux qui liront ces lignes </p>
        <gap/> 
        <p>Votre Saint Nom en soit béni, loué et glorifié. Amen.</p></doc>
        """
wien_xml_raw = """
    <doc>
    <line>louerent Dieu, mais mr. le C. de Sli. ne voulut  pas appliquer mon voyage a mon avantage 
    il crut que cela rasembloit fort a l’avanture, et que la peur de me confesser au RP. Br.
    m’avoit fait aller a Hitzing Je ne m’excuse pas sur ce point, 
    laissant au jugement de ceux qui liront ces ligne votre st nom en soitloué, et glorifié, amen.</line>
    </doc>
"""

方案1失效原因分析

方案1基于BeautifulSoup实现,核心问题有两点:

  • 匹配对象错误:代码中拿整个文本节点的内容和单个人名做token_sort_ratio对比,而人名只是文本节点中的一小部分(比如<line>标签包含大段文本),两者的相似度远低于设定的90阈值,导致条件永远不触发。
  • 替换目标错误:就算触发匹配,代码尝试替换的是第一份中的原人名(如mr. l Conte de Sle.),但第二份中的人名是拼写变体(如mr. le C. de Sli.),原字符串根本不存在,替换操作无效。

方案1修正代码

from bs4 import BeautifulSoup, Tag
from fuzzywuzzy import fuzz, process

# 解析第一份文档,提取人名
soup1 = BeautifulSoup(aix_xml_raw, 'xml')
pers_names = [tag.text.strip() for tag in soup1.find_all('persName')]

# 解析第二份文档
soup2 = BeautifulSoup(wien_xml_raw, 'xml')

# 遍历所有文本节点处理
for node in soup2.find_all(text=True):
    original_text = node.strip()
    if not original_text:
        continue  # 跳过空白节点
    
    new_text = original_text
    for name in pers_names:
        # 用partial_ratio匹配文本中与人名相似的子串
        match, score = process.extractOne(name, [original_text], scorer=fuzz.partial_ratio)
        if score > 80:
            # 创建persName标签,替换找到的变体子串
            new_tag = Tag(name='persName')
            new_tag.string = match
            new_text = new_text.replace(match, str(new_tag))
    
    # 替换原节点内容
    if new_text != original_text:
        node.replace_with(new_text)

# 输出结果
print(soup2.prettify())

方案2失效原因分析

方案2基于xml.etree.ElementTree和difflib实现,核心问题和方案1一致:

  • 匹配对象错误:拿整个<line>标签的文本内容和单个人名做相似度对比,长文本与短人名的相似度远低于0.8的阈值,条件无法触发。
  • 替换目标错误:尝试替换原人名而非第二份中的拼写变体,导致替换操作找不到目标字符串。

方案2修正代码

import xml.etree.ElementTree as ET
from difflib import SequenceMatcher

def get_person_names(xml_str):
    person_names = []
    root = ET.fromstring(xml_str)
    for pers_name in root.iter('persName'):
        person_names.append(pers_name.text.strip())
    return person_names

def find_best_match(s, target, threshold=0.8):
    """在字符串s中找到与target最相似的子串"""
    max_ratio = 0
    best_match = ""
    len_target = len(target)
    len_s = len(s)
    
    # 滑动窗口遍历与目标长度相近的子串
    for delta in [-2, -1, 0, 1, 2]:
        window_len = len_target + delta
        if window_len <= 0 or window_len > len_s:
            continue
        for i in range(len_s - window_len + 1):
            window = s[i:i+window_len]
            ratio = SequenceMatcher(None, target.lower(), window.lower()).ratio()
            if ratio > max_ratio:
                max_ratio = ratio
                best_match = window
    
    return (best_match, max_ratio) if max_ratio >= threshold else (None, 0)

def tag_person_names(xml_str, person_names):
    root = ET.fromstring(xml_str)
    for line in root.iter('line'):
        if not line.text:
            continue
        original_text = line.text
        tagged_text = original_text
        
        for name in person_names:
            match, _ = find_best_match(original_text, name)
            if match:
                tagged_text = tagged_text.replace(match, f'<persName>{match}</persName>')
        
        line.text = tagged_text
    return ET.tostring(root, encoding='unicode')

# 执行流程
person_names = get_person_names(aix_xml_raw)
tagged_wien_xml_raw = tag_person_names(wien_xml_raw, person_names)
print(tagged_wien_xml_raw)

内容的提问来源于stack exchange,提问作者Selina Galka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 09:25:08