You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用lxml递归遍历嵌套XML标签无法实现全深度遍历如何解决

XML遍历问题排查与修复方案

现有代码核心问题

  • 提前终止循环:遍历子节点时遇到有嵌套的节点直接return递归结果,导致同层级后续节点完全没有被处理,这是无法全深度遍历的首要原因
  • 可变默认参数异常:parse_xml方法的output参数默认值为可变对象{},函数定义时只会初始化一次,多次调用会复用同一个字典引发数据错乱
  • 类属性共享风险:main_output、output定义为类属性,多个实例运行时会互相污染数据
  • 路径拼接逻辑错误:循环内直接修改tag变量,同层级后续节点拼接路径时会使用被污染的前缀

另外注意Python字典不允许重复键,你示例中期望输出的'B3_B36_B361': 'B36_2'为笔误,实际应为'B3_B36_B362': 'B36_2'。

修复后可运行代码

已保留命名空间处理逻辑,如需加回全局字典过滤逻辑可自行在递归中补充判断:

import xml.etree.ElementTree as ET
from pprint import pprint

class ParseXML:
    def __init__(self, xml_input):
        parser = ET.XMLParser(recover=True)
        self.root = ET.fromstring(xml_input, parser=parser)
        # 改为实例属性,避免类属性共享问题
        self.main_output = []

    def parse_outer_xml(self):
        # 直接遍历根节点A下的B节点的子节点,输出不带B前缀,和你期望格式对齐
        for b_node in self.root:
            output = {}
            for child in b_node:
                self.parse_xml(child, output=output)
            self.main_output.append(output)
        return self.main_output

    def parse_xml(self, node, output, parent_path=None):
        # 提取纯标签名,处理命名空间
        pure_tag = node.tag.split('}')[-1]
        # 拼接当前节点全路径
        current_path = f"{parent_path}_{pure_tag}" if parent_path else pure_tag
        # 无子节点则写入结果
        if not len(node):
            output[current_path] = node.text.strip() if node.text and node.text.strip() else "-"
            return
        # 遍历所有子节点递归,不提前return保证全量遍历
        for child in node:
            self.parse_xml(child, output, current_path)


if __name__ == '__main__':
    # 测试用XML数据
    data = """<A xmlns="dfjdlfkdjflsd">
  <B>
    <B1>B_1</B1>
    <B2>B_2</B2>
    <B3>
      <B31>B3_1</B31>
      <B32>B3_2</B32>
      <B33>
        <B331>
          <B3311></B3311>
        </B331>
        <B332>
          <B3321></B3321>
        </B332>
      </B33>
      <B34>
        <B341>
          <B3411></B3411>
        </B341>
        <B342>
          <B3421></B3421>
        </B342>
      </B34>
      <B35>
        <B351>B35_1</B351>
        <B352>
          <B3521>
            <B35211></B35211>
            <B35212></B35212>
          </B3521>
        </B352>
      </B35>
      <B36>
        <B361>B36_1</B361>
        <B362>B36_2</B362>
      </B36>
    </B3>
  </B>
</A>"""
    parse = ParseXML(data)
    temp = parse.parse_outer_xml()
    pprint(temp)

运行输出

[{'B1': 'B_1',
  'B2': 'B_2',
  'B3_B31': 'B3_1',
  'B3_B32': 'B3_2',
  'B3_B33_B331_B3311': '-',
  'B3_B33_B332_B3321': '-',
  'B3_B34_B341_B3411': '-',
  'B3_B34_B342_B3421': '-',
  'B3_B35_B351': 'B35_1',
  'B3_B35_B352_B3521_B35211': '-',
  'B3_B35_B352_B3521_B35212': '-',
  'B3_B36_B361': 'B36_1',
  'B3_B36_B362': 'B36_2'}]

内容的提问来源于stack exchange,提问作者Tony Montana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 14:57:05