如何在Python中使用ElementTree获取XML指定文本内容?
问题
使用ElementTree遍历XML标签时,<result1>标签的text返回None,而非预期的“This is the Result of Chrome Browser”。
XML示例
<?xml version="1.0" encoding="UTF-8"?> <article> <result1 id="val1"><h2>Google Chrome</h2> This is the Result of Chrome Browser<p>Google Chrome is a web browser developed by Google, released in 2008. Chrome is the world's most popular web browser today!</p></result1> </article>
测试代码
import xml.etree.ElementTree as ET treexml = ET.parse('exam.xml') for elemintree in treexml.iter(): print(elemintree.tag,elemintree.text)
当前输出
article result1 None h2 Google Chrome p Google Chrome is a web browser developed by Google, released in 2008. Chrome is the world's most popular web browser today!
期望输出
h2 Google Chrome This is the Result of Chrome Browser p Google Chrome is a web browser developed by Google, released in 2008. Chrome is the world's most popular web browser today!
解决方法
原因分析
xml.etree.ElementTree中,元素的text属性仅存储该元素第一个子元素之前的文本。<result1>的第一个子元素是<h2>,所以result1.text为None;你需要的目标文本实际是<h2>元素的tail属性(即子元素结束后到下一个子元素开始前的内容)。
修正代码
要获取所有需要的内容,需同时遍历元素的text和tail属性,并过滤空白内容:
import xml.etree.ElementTree as ET treexml = ET.parse('exam.xml') for elem in treexml.iter(): # 输出元素标签与非空text内容 if elem.text and elem.text.strip(): print(f"{elem.tag} {elem.text.strip()}") # 输出元素的非空tail文本 if elem.tail and elem.tail.strip(): print(elem.tail.strip())
输出结果
h2 Google Chrome This is the Result of Chrome Browser p Google Chrome is a web browser developed by Google, released in 2008. Chrome is the world's most popular web browser today!
关键说明
tail属性专门用于存储元素之后、父元素下一个子元素之前的文本内容。- 使用
strip()是为了去除文本前后的换行和空白字符,避免输出多余的空行。
内容的提问来源于stack exchange,提问作者priyanka manogaran
相关产品推荐
相关产品推荐

