You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用ElementTree删除XML文件中的空白与空列?

用ElementTree清理XML中的多余空白与空节点

一、清理文本节点的多余空白

XML中的多余空白常存在于元素的text或tail属性中(比如文本前后的空格、换行符)。你可以遍历所有元素,对这些属性做去空白处理:

import xml.etree.ElementTree as ET

def clean_whitespace(element):
    # 处理当前元素的文本内容
    if element.text is not None:
        stripped_text = element.text.strip()
        element.text = stripped_text if stripped_text else None
    # 处理元素闭合标签后的尾随文本
    if element.tail is not None:
        stripped_tail = element.tail.strip()
        element.tail = stripped_tail if stripped_tail else None
    # 递归处理子元素
    for child in list(element):
        clean_whitespace(child)

二、删除空元素(空列)

这里的“空列”指无任何子元素、且text和tail都为空的元素。可以通过递归遍历,移除符合条件的空节点:

def remove_empty_elements(element):
    # 从后往前遍历子元素,避免删除操作打乱遍历顺序
    for child in reversed(list(element)):
        remove_empty_elements(child)
        # 判断是否为完全空的元素
        if not list(child) and not child.text and not child.tail:
            element.remove(child)

三、结合现有代码的完整示例

把上述清理逻辑和你已有的节点删除代码结合,完整流程如下:

import xml.etree.ElementTree as ET

def clean_whitespace(element):
    if element.text is not None:
        stripped_text = element.text.strip()
        element.text = stripped_text if stripped_text else None
    if element.tail is not None:
        stripped_tail = element.tail.strip()
        element.tail = stripped_tail if stripped_tail else None
    for child in list(element):
        clean_whitespace(child)

def remove_empty_elements(element):
    for child in reversed(list(element)):
        remove_empty_elements(child)
        if not list(child) and not child.text and not child.tail:
            element.remove(child)

# 解析XML文件
tree = ET.parse('SampleData.xml')
root = tree.getroot()

# 1. 清理所有多余空白
clean_whitespace(root)

# 2. 删除空元素
remove_empty_elements(root)

# 3. 执行你原有的节点删除逻辑
for country in root.findall('country'):
    description_node = country.find('description')
    # 先判断节点是否存在,避免报错
    if description_node is not None and description_node.text == 'Liechtenstein has a lot of flowers.':
        country.remove(description_node)

# 保存处理后的XML文件
tree.write('CleanedSampleData.xml', encoding='utf-8', xml_declaration=True)

说明

  • clean_whitespace函数会递归处理所有层级的元素,去掉文本前后的空白,处理后为空的内容会设为None,避免残留空字符串。
  • remove_empty_elements函数从后往前遍历子元素,防止删除节点后导致的遍历索引混乱,确保所有空元素都被清理。

内容的提问来源于stack exchange,提问作者joemamah24

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 19:35:44