You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中使用ElementTree获取XML指定文本内容?

问题

使用ElementTree遍历XML标签时,<result1>标签的text返回None,而非预期的“This is the Result of Chrome Browser”。

XML示例

<?xml version="1.0" encoding="UTF-8"?>
<article>
 <result1 id="val1"><h2>Google Chrome</h2>
 This is the Result of Chrome Browser<p>Google Chrome is a web browser developed by Google, released in 2008. Chrome is the world's 
  most popular web browser today!</p></result1>
</article>

测试代码

import xml.etree.ElementTree as ET
treexml = ET.parse('exam.xml')
for elemintree in treexml.iter():
    print(elemintree.tag,elemintree.text)

当前输出

article
result1 None
h2 Google Chrome
p Google Chrome is a web browser developed by Google, released in 2008. Chrome is the world's 
  most popular web browser today!

期望输出

h2 Google Chrome
This is the Result of Chrome Browser
p Google Chrome is a web browser developed by Google, released in 2008. Chrome is the world's 
  most popular web browser today!

解决方法

原因分析

xml.etree.ElementTree中,元素的text属性仅存储该元素第一个子元素之前的文本。<result1>的第一个子元素是<h2>,所以result1.text为None;你需要的目标文本实际是<h2>元素的tail属性(即子元素结束后到下一个子元素开始前的内容)。

修正代码

要获取所有需要的内容,需同时遍历元素的text和tail属性,并过滤空白内容:

import xml.etree.ElementTree as ET

treexml = ET.parse('exam.xml')
for elem in treexml.iter():
    # 输出元素标签与非空text内容
    if elem.text and elem.text.strip():
        print(f"{elem.tag} {elem.text.strip()}")
    # 输出元素的非空tail文本
    if elem.tail and elem.tail.strip():
        print(elem.tail.strip())

输出结果

h2 Google Chrome
This is the Result of Chrome Browser
p Google Chrome is a web browser developed by Google, released in 2008. Chrome is the world's most popular web browser today!

关键说明

  • tail属性专门用于存储元素之后、父元素下一个子元素之前的文本内容。
  • 使用strip()是为了去除文本前后的换行和空白字符,避免输出多余的空行。

内容的提问来源于stack exchange,提问作者priyanka manogaran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 09:47:28