You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3.6下按父节点拆分XML标签列表的技术求助

按父节点拆分XML标签列表的解决方案

我来帮你解决这个按父节点拆分标签的问题!在Python3.6里,用自带的xml.etree.ElementTree库就能轻松搞定,下面给你具体的步骤和代码示例:

1. 先看一个示例XML结构

假设你的XML文件大致是这样的(你可以替换成自己的实际结构):

<root>
    <discover>
        <product>Apple</product>
        <price>999</price>
        <stock>100</stock>
    </discover>
    <discover>
        <product>Banana</product>
        <price>10</price>
    </discover>
    <other_node>
        <category>Fruit</category>
    </other_node>
</root>

2. 基础实现:拆分每个discover的直接子标签

下面的代码会遍历所有discover父节点,把每个节点下的直接子标签单独存成列表:

import xml.etree.ElementTree as ET

# 解析你的XML文件(替换成实际文件路径)
tree = ET.parse('your_xml_file.xml')
root = tree.getroot()

# 存储每个discover节点的标签列表
discover_tag_groups = []

# 遍历所有discover节点(不管在XML的哪个层级)
for discover_node in root.findall('.//discover'):
    # 获取当前discover下所有直接子节点的标签名
    child_tags = [child.tag for child in discover_node]
    discover_tag_groups.append(child_tags)

# 打印结果
for index, tags in enumerate(discover_tag_groups, start=1):
    print(f"第{index}个discover节点的标签列表: {tags}")

运行这段代码后,输出会是:

第1个discover节点的标签列表: ['product', 'price', 'stock']
第2个discover节点的标签列表: ['product', 'price']

3. 扩展:获取子节点的完整信息(标签+内容)

如果你不仅需要标签名,还想拿到标签对应的文本内容,可以修改代码:

import xml.etree.ElementTree as ET

tree = ET.parse('your_xml_file.xml')
root = tree.getroot()

discover_child_info = []
for discover_node in root.findall('.//discover'):
    child_details = [
        {
            'tag': child.tag,
            'content': child.text.strip() if child.text else None
        }
        for child in discover_node
    ]
    discover_child_info.append(child_details)

# 打印结果
for index, details in enumerate(discover_child_info, start=1):
    print(f"\n第{index}个discover节点的子节点详情:")
    for item in details:
        print(f"- {item['tag']}: {item['content']}")

4. 进阶:递归获取所有后代标签(包括嵌套子节点)

如果你的discover节点下还有多层嵌套的子节点,比如<discover><info><name>xxx</name></info></discover>,可以用递归函数获取所有后代标签:

import xml.etree.ElementTree as ET

def get_all_descendant_tags(node):
    """递归获取节点的所有后代标签名"""
    tags = [node.tag]
    for child in node:
        tags.extend(get_all_descendant_tags(child))
    return tags

tree = ET.parse('your_xml_file.xml')
root = tree.getroot()

discover_all_tags = []
for discover_node in root.findall('.//discover'):
    all_tags = get_all_descendant_tags(discover_node)
    discover_all_tags.append(all_tags)

# 打印结果
for index, tags in enumerate(discover_all_tags, start=1):
    print(f"第{index}个discover节点的所有标签(含嵌套): {tags}")

注意事项

  • 请把代码中的your_xml_file.xml替换成你实际的XML文件路径;
  • 如果你的XML是字符串形式(不是文件),可以用ET.fromstring(xml_string)来解析,而不是ET.parse();
  • Python3.6中Element对象的getchildren()方法已被弃用,直接用for child in node遍历子节点即可。

内容的提问来源于stack exchange,提问作者ilapasle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:24:04