如何用Python在XML文档中添加<person>标签提取方括号内人名
解决XML提取人名并插入
<person>标签的问题 你的代码已经完成了人名提取,但在标签创建和插入时存在两个核心问题:
- 只处理了每个
newsDocument下的第一个newsFrom节点,忽略了同文档内的其他newsFrom - 没有准确定位
<transc>标签的位置,导致插入的标签位置混乱
以下是修正后的完整代码:
import xml.etree.ElementTree as ET import re # 解析XML文件 tree = ET.parse('1649.xml') root = tree.getroot() # 遍历所有newsFrom节点(覆盖所有文档下的所有来源) for news_from in root.findall(".//newsFrom"): # 获取当前newsFrom下的transc标签 transc_elem = news_from.find("transc") if transc_elem is None or transc_elem.text is None: continue # 跳过无文本的transc标签 # 提取方括号内的人名 people = re.findall(r"\[(.*?)\]", transc_elem.text) # 找到transc在父节点中的索引位置,确保插入到它的下方 transc_index = list(news_from).index(transc_elem) # 为每个人名创建person标签并插入到transc之后 for name in people: person_elem = ET.Element("person") person_elem.text = name # 插入到transc的下一个位置,保持顺序 news_from.insert(transc_index + 1, person_elem) # 每插入一个标签,后续标签的插入位置要+1 transc_index += 1 # 保存修改后的XML文件 tree.write('modified_1649.xml', encoding='utf-8', xml_declaration=True)
关键修复点说明
- 遍历所有
newsFrom:使用".//newsFrom"XPath表达式直接定位所有层级下的newsFrom节点,避免遗漏 - 空值判断:增加对
transc标签和其文本的非空检查,防止因缺失内容导致的报错 - 精准插入位置:通过
list(news_from).index(transc_elem)获取transc在父节点中的位置,然后依次插入<person>标签到它的下一个位置,确保结构和示例一致 - 动态调整插入索引:每插入一个
<person>标签后,索引值加1,保证后续标签依次排列在transc下方
修改后的XML效果示例
<newsFrom> <from date="15/01/1649" dateUnsure="y">London</from> <transc>Questo Parlamento generale Farfax [Thomas Fairfax, 3rd Lord Fairfax of Cameron] et suo consiglio dio et ordinato di pocessare il re [Charles I, King of England]</transc> <person>Thomas Fairfax, 3rd Lord Fairfax of Cameron</person> <person>Charles I, King of England</person> <newsTopic>Military</newsTopic> <wordCount>103</wordCount> <position>1</position> </newsFrom>
内容的提问来源于stack exchange,提问作者MessyCoder
相关产品推荐
相关产品推荐

