You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用LXML移除XML文件中剥离标签后遗留的空行?

移除XML标签剥离后遗留空行的解决办法

问题情况

输入的in.xml内容是:

<?xml version="1.0" encoding="ASCII"?>
<a>
  <b>
    <c>abc</c>
  </b>
</a>

用这段Python代码剥离<b>标签并输出到out.xml:

from lxml import etree

tree = etree.parse('in.xml', parser=etree.XMLParser())
root = tree.getroot()
etree.strip_tags(root, 'b')
tree.write('out.xml', xml_declaration=True)

结果生成的out.xml里多了不少空行:

<?xml version='1.0' encoding='ASCII'?>
<a>
  
    <c>abc</c>
  
</a>

两种解决方法

方法1:解析时直接忽略空白文本

初始化XML解析器的时候加上remove_blank_text=True,解析阶段就会自动清理没用的空白节点,剥离标签后自然不会留空行:

from lxml import etree

# 开启空白文本移除功能
tree = etree.parse('in.xml', parser=etree.XMLParser(remove_blank_text=True))
root = tree.getroot()
etree.strip_tags(root, 'b')
# 写入时指定编码,同时开启格式化让输出更整洁
tree.write('out.xml', xml_declaration=True, encoding='ASCII', pretty_print=True)

生成的out.xml就正常了:

<?xml version='1.0' encoding='ASCII'?>
<a>
    <c>abc</c>
</a>

方法2:手动清理残留的空白节点

如果不想改解析器参数,也可以在剥离标签后,手动遍历节点删掉纯空白的文本节点:

from lxml import etree

tree = etree.parse('in.xml', parser=etree.XMLParser())
root = tree.getroot()
etree.strip_tags(root, 'b')

# 遍历根节点的子元素,把纯空白的文本节点删掉
for child in list(root):
    if isinstance(child, etree._ElementUnicodeResult) and child.strip() == '':
        root.remove(child)

tree.write('out.xml', xml_declaration=True, encoding='ASCII')

这样也能彻底去掉那些多余的空行。

内容的提问来源于stack exchange,提问作者mrgou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 17:40:28