You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3.x中xml.etree.ElementTree无法获取元素后续文本求助

问题:ElementTree解析XML时无法获取子标签后的文本

在Windows环境下用Python 3.12的xml.etree.ElementTree解析testfile.xml时,第三个<hl2>元素的text属性返回None,拿不到该元素内<em>标签后的文本“e are always calling each other names.”。

XML文件内容

<?xml version="1.0" encoding="utf-8"?>
<body>
  <body.head>
    <hedline>
      <hl1 style="header">All the things we lost that summer</hl1>
      <hl2 style="standfirst">It was the promise of seals that sold Virginia on this mission.</hl2>
      <hl2 style="dropcap-large"><em class="dropcap">W</em>e are always calling each other names.</hl2>
    </hedline>
  </body.head>
</body>

原解析脚本

import xml.etree.ElementTree as ET
tree = ET.parse('testfile.xml')
root = tree.getroot()
if root.find('body.head') is not None:
    if root.find('body.head').find('hedline') is not None:
        for child1 in root.find('body.head').find('hedline'):
            print("Tag    level 1:" + child1.tag)
            print("Attrib level 1:" + str(child1.attrib))
            print("Text   level 1:" + str(child1.text) + "\n")
            for child2 in child1:
                print("Tag    level 2:" + child2.tag)
                print("Attrib level 2:" + str(child2.attrib))
                print("Text   level 2:" + str(child2.text))

原执行结果

Tag    level 1:hl1
Attrib level 1:{'style': 'header'}
Text   level 1:All the things we lost that summer

Tag    level 1:hl2
Attrib level 1:{'style': 'standfirst'}
Text   level 1:It was the promise of seals that sold Virginia on this mission.

Tag    level 1:hl2
Attrib level 1:{'style': 'dropcap-large'}
Text   level 1:None  <-- THIS IS THE PROBLEM

Tag    level 2:em
Attrib level 2:{'class': 'dropcap'}
Text   level 2:W

原因分析

ElementTree中,元素的text属性仅对应该元素开始标签与第一个子元素之间的文本。第三个<hl2>的第一个内容是<em>子标签,所以hl2.text为None;而<em>标签之后的文本,实际存储在<em>元素的tail属性中。

解决方法

方法1:用ET.tostring()直接提取完整文本

这是最简便的方式,能直接获取元素下所有文本内容,无需手动拼接:

import xml.etree.ElementTree as ET
tree = ET.parse('testfile.xml')
root = tree.getroot()

# 简化节点查找
hedline = root.find('body.head/hedline')
if hedline:
    for child1 in hedline:
        print("Tag    level 1:" + child1.tag)
        print("Attrib level 1:" + str(child1.attrib))
        # 提取当前元素下的所有文本,encoding='unicode'返回字符串
        full_text = ET.tostring(child1, encoding='unicode', method='text').strip()
        print("Full Text:" + full_text + "\n")
        for child2 in child1:
            print("Tag    level 2:" + child2.tag)
            print("Attrib level 2:" + str(child2.attrib))
            print("Text   level 2:" + str(child2.text))
            print("Tail   level 2:" + str(child2.tail) + "\n")

执行后,第三个<hl2>的Full Text会显示完整的"We are always calling each other names.",同时能看到<em>的tail属性值就是缺失的文本部分。

方法2:手动拼接text与子元素的tail

如果需要更精细的文本控制,可以手动拼接各部分内容:

import xml.etree.ElementTree as ET
tree = ET.parse('testfile.xml')
root = tree.getroot()

hedline = root.find('body.head/hedline')
if hedline:
    for child1 in hedline:
        print("Tag    level 1:" + child1.tag)
        print("Attrib level 1:" + str(child1.attrib))
        
        # 收集所有文本部分
        text_parts = []
        if child1.text:
            text_parts.append(child1.text)
        for sub_elem in child1:
            text_parts.append(sub_elem.text)
            if sub_elem.tail:
                text_parts.append(sub_elem.tail)
        
        full_text = ''.join(text_parts).strip()
        print("Full Text:" + full_text + "\n")

这个方法会依次拼接父元素的text、每个子元素的text和tail,最终得到完整文本。


内容的提问来源于stack exchange,提问作者Martijn Nabben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 07:59:52