基于Python模糊匹配为XML标记persName标签的失效问题排查
模糊匹配无法生效的原因及修复方案
问题背景
我有两份XML文件:第一份aix_xml_raw中的人名已用<persName>标签完成标记;第二份wien_xml_raw内容相近,但存在拼写差异且新增了部分段落。我希望通过模糊匹配(例如第一份中的mr. l Conte de Sle匹配第二份的mr. le C. de Sli.),将第一份中<persName>元素的内容在第二份中定位并标记,但添加模糊匹配的判断条件后代码无法生效,移除该条件后反而能正常运行,请问这是为什么?
原始XML示例
aix_xml_raw = """ <doc><p>Louërent Dieu, mais <persName>mr. l C<ex>onte</ex> de Sle.</persName> ne voulut pas appliquer mon voyage a mon avantage il crut que cela ressembloit fort a l’avanture, et que la peur de me confesser au <persName>RP. Br.</persName> m’avoit fait aller a <placeName>Hilzing</placeName> Je ne m’excuse point, laissant au jugement de ceux qui liront ces lignes </p> <gap/> <p>Votre Saint Nom en soit béni, loué et glorifié. Amen.</p></doc> """ wien_xml_raw = """ <doc> <line>louerent Dieu, mais mr. le C. de Sli. ne voulut pas appliquer mon voyage a mon avantage il crut que cela rasembloit fort a l’avanture, et que la peur de me confesser au RP. Br. m’avoit fait aller a Hitzing Je ne m’excuse pas sur ce point, laissant au jugement de ceux qui liront ces ligne votre st nom en soitloué, et glorifié, amen.</line> </doc> """
方案1失效原因分析
方案1基于BeautifulSoup实现,核心问题有两点:
- 匹配对象错误:代码中拿整个文本节点的内容和单个人名做
token_sort_ratio对比,而人名只是文本节点中的一小部分(比如<line>标签包含大段文本),两者的相似度远低于设定的90阈值,导致条件永远不触发。 - 替换目标错误:就算触发匹配,代码尝试替换的是第一份中的原人名(如
mr. l Conte de Sle.),但第二份中的人名是拼写变体(如mr. le C. de Sli.),原字符串根本不存在,替换操作无效。
方案1修正代码
from bs4 import BeautifulSoup, Tag from fuzzywuzzy import fuzz, process # 解析第一份文档,提取人名 soup1 = BeautifulSoup(aix_xml_raw, 'xml') pers_names = [tag.text.strip() for tag in soup1.find_all('persName')] # 解析第二份文档 soup2 = BeautifulSoup(wien_xml_raw, 'xml') # 遍历所有文本节点处理 for node in soup2.find_all(text=True): original_text = node.strip() if not original_text: continue # 跳过空白节点 new_text = original_text for name in pers_names: # 用partial_ratio匹配文本中与人名相似的子串 match, score = process.extractOne(name, [original_text], scorer=fuzz.partial_ratio) if score > 80: # 创建persName标签,替换找到的变体子串 new_tag = Tag(name='persName') new_tag.string = match new_text = new_text.replace(match, str(new_tag)) # 替换原节点内容 if new_text != original_text: node.replace_with(new_text) # 输出结果 print(soup2.prettify())
方案2失效原因分析
方案2基于xml.etree.ElementTree和difflib实现,核心问题和方案1一致:
- 匹配对象错误:拿整个
<line>标签的文本内容和单个人名做相似度对比,长文本与短人名的相似度远低于0.8的阈值,条件无法触发。 - 替换目标错误:尝试替换原人名而非第二份中的拼写变体,导致替换操作找不到目标字符串。
方案2修正代码
import xml.etree.ElementTree as ET from difflib import SequenceMatcher def get_person_names(xml_str): person_names = [] root = ET.fromstring(xml_str) for pers_name in root.iter('persName'): person_names.append(pers_name.text.strip()) return person_names def find_best_match(s, target, threshold=0.8): """在字符串s中找到与target最相似的子串""" max_ratio = 0 best_match = "" len_target = len(target) len_s = len(s) # 滑动窗口遍历与目标长度相近的子串 for delta in [-2, -1, 0, 1, 2]: window_len = len_target + delta if window_len <= 0 or window_len > len_s: continue for i in range(len_s - window_len + 1): window = s[i:i+window_len] ratio = SequenceMatcher(None, target.lower(), window.lower()).ratio() if ratio > max_ratio: max_ratio = ratio best_match = window return (best_match, max_ratio) if max_ratio >= threshold else (None, 0) def tag_person_names(xml_str, person_names): root = ET.fromstring(xml_str) for line in root.iter('line'): if not line.text: continue original_text = line.text tagged_text = original_text for name in person_names: match, _ = find_best_match(original_text, name) if match: tagged_text = tagged_text.replace(match, f'<persName>{match}</persName>') line.text = tagged_text return ET.tostring(root, encoding='unicode') # 执行流程 person_names = get_person_names(aix_xml_raw) tagged_wien_xml_raw = tag_person_names(wien_xml_raw, person_names) print(tagged_wien_xml_raw)
内容的提问来源于stack exchange,提问作者Selina Galka
相关产品推荐
相关产品推荐

