Python 3.x中xml.etree.ElementTree无法获取元素后续文本求助
问题:ElementTree解析XML时无法获取子标签后的文本
在Windows环境下用Python 3.12的xml.etree.ElementTree解析testfile.xml时,第三个<hl2>元素的text属性返回None,拿不到该元素内<em>标签后的文本“e are always calling each other names.”。
XML文件内容
<?xml version="1.0" encoding="utf-8"?> <body> <body.head> <hedline> <hl1 style="header">All the things we lost that summer</hl1> <hl2 style="standfirst">It was the promise of seals that sold Virginia on this mission.</hl2> <hl2 style="dropcap-large"><em class="dropcap">W</em>e are always calling each other names.</hl2> </hedline> </body.head> </body>
原解析脚本
import xml.etree.ElementTree as ET tree = ET.parse('testfile.xml') root = tree.getroot() if root.find('body.head') is not None: if root.find('body.head').find('hedline') is not None: for child1 in root.find('body.head').find('hedline'): print("Tag level 1:" + child1.tag) print("Attrib level 1:" + str(child1.attrib)) print("Text level 1:" + str(child1.text) + "\n") for child2 in child1: print("Tag level 2:" + child2.tag) print("Attrib level 2:" + str(child2.attrib)) print("Text level 2:" + str(child2.text))
原执行结果
Tag level 1:hl1 Attrib level 1:{'style': 'header'} Text level 1:All the things we lost that summer Tag level 1:hl2 Attrib level 1:{'style': 'standfirst'} Text level 1:It was the promise of seals that sold Virginia on this mission. Tag level 1:hl2 Attrib level 1:{'style': 'dropcap-large'} Text level 1:None <-- THIS IS THE PROBLEM Tag level 2:em Attrib level 2:{'class': 'dropcap'} Text level 2:W
原因分析
ElementTree中,元素的text属性仅对应该元素开始标签与第一个子元素之间的文本。第三个<hl2>的第一个内容是<em>子标签,所以hl2.text为None;而<em>标签之后的文本,实际存储在<em>元素的tail属性中。
解决方法
方法1:用ET.tostring()直接提取完整文本
这是最简便的方式,能直接获取元素下所有文本内容,无需手动拼接:
import xml.etree.ElementTree as ET tree = ET.parse('testfile.xml') root = tree.getroot() # 简化节点查找 hedline = root.find('body.head/hedline') if hedline: for child1 in hedline: print("Tag level 1:" + child1.tag) print("Attrib level 1:" + str(child1.attrib)) # 提取当前元素下的所有文本,encoding='unicode'返回字符串 full_text = ET.tostring(child1, encoding='unicode', method='text').strip() print("Full Text:" + full_text + "\n") for child2 in child1: print("Tag level 2:" + child2.tag) print("Attrib level 2:" + str(child2.attrib)) print("Text level 2:" + str(child2.text)) print("Tail level 2:" + str(child2.tail) + "\n")
执行后,第三个<hl2>的Full Text会显示完整的"We are always calling each other names.",同时能看到<em>的tail属性值就是缺失的文本部分。
方法2:手动拼接text与子元素的tail
如果需要更精细的文本控制,可以手动拼接各部分内容:
import xml.etree.ElementTree as ET tree = ET.parse('testfile.xml') root = tree.getroot() hedline = root.find('body.head/hedline') if hedline: for child1 in hedline: print("Tag level 1:" + child1.tag) print("Attrib level 1:" + str(child1.attrib)) # 收集所有文本部分 text_parts = [] if child1.text: text_parts.append(child1.text) for sub_elem in child1: text_parts.append(sub_elem.text) if sub_elem.tail: text_parts.append(sub_elem.tail) full_text = ''.join(text_parts).strip() print("Full Text:" + full_text + "\n")
这个方法会依次拼接父元素的text、每个子元素的text和tail,最终得到完整文本。
内容的提问来源于stack exchange,提问作者Martijn Nabben
相关产品推荐
相关产品推荐

