You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用ElementTree提取XML管理者信息?findall无结果原因解析

问题描述

尝试用ElementTree提取XML响应数据,目标XML文件xmlresponse.xml结构如下:

<result xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="https://somewhere.co.uk/">
    <count>1</count>
    <pageInformation>
        <offset>0</offset>
        <size>10</size>
    </pageInformation>
    <items>
        <person uuid="1">
            <name>
                <firstName>John</firstName>
                <lastName>Doe</lastName>
            </name>
            <ManagedByRelations>
                <managedByRelation Id="1234">
                    <manager uuid="2">
                        <name formatted="false">
                            <text>Jane Doe</text>
                        </name>
                    </manager>
                    <managementPercentage>30</managementPercentage>
                    <period>
                        <startDate>2019-09-26</startDate>
                    </period>

                </managedByRelation>
                <managedByRelation Id="1234">
                    <manager uuid="3">
                        <name formatted="false">
                            <text>Joe Bloggs</text>
                        </name>
                    </manager>
                    <managementPercentage>70</managementPercentage>
                    <period>
                        <startDate>2019-09-26</startDate>
                    </period>
                </managedByRelation>
            </ManagedByRelations>
            <fte>0.0</fte>
        </person>
    </items>
</result>

希望提取管理者的姓名、ID及开始日期列表,但执行以下代码时,root.findall('managedByRelation')未返回任何结果:

from xml.etree.ElementTree import Element, ParseError, fromstring, tostring, parse

tree = parse('xmlresponse.xml')
root = tree.getroot()

for manager in root.findall('managedByRelation'):
    print(manager)

已知可以通过list(root.iter())遍历整个XML树,但想了解findall方法未按预期工作的原因,同时获取正确提取目标信息的方法。

问题分析与解决

为什么findall('managedByRelation')无结果

findall()默认只查找当前节点的直接子节点,而managedByRelation的层级是result -> items -> person -> ManagedByRelations -> managedByRelation,它不是根节点(<result>)的直接子节点,所以直接调用root.findall('managedByRelation')找不到匹配项。

另外,findall()支持XPath表达式,但默认路径是相对当前节点的,必须写出完整相对路径或用//前缀匹配任意层级节点。

正确提取目标信息的方法

方式1:使用完整相对XPath路径

明确写出从根节点到目标节点的路径:

from xml.etree.ElementTree import parse

tree = parse('xmlresponse.xml')
root = tree.getroot()

# 匹配指定路径下的所有managedByRelation节点
for relation in root.findall('./items/person/ManagedByRelations/managedByRelation'):
    # 提取managedByRelation的Id属性
    relation_id = relation.get('Id')
    # 提取管理者姓名
    manager_name = relation.find('./manager/name/text').text
    # 提取开始日期
    start_date = relation.find('./period/startDate').text
    
    print(f"管理者ID: {relation_id}, 姓名: {manager_name}, 开始日期: {start_date}")

方式2:使用//匹配任意层级节点

如果不想写完整路径,用//可以匹配任意深度的目标节点:

from xml.etree.ElementTree import parse

tree = parse('xmlresponse.xml')
root = tree.getroot()

# 匹配所有层级下的managedByRelation节点
for relation in root.findall('.//managedByRelation'):
    relation_id = relation.get('Id')
    manager_name = relation.find('manager/name/text').text
    start_date = relation.find('period/startDate').text
    
    print(f"管理者ID: {relation_id}, 姓名: {manager_name}, 开始日期: {start_date}")

输出结果

两种方法都会输出:

管理者ID: 1234, 姓名: Jane Doe, 开始日期: 2019-09-26
管理者ID: 1234, 姓名: Joe Bloggs, 开始日期: 2019-09-26

内容的提问来源于stack exchange,提问作者abinitio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 00:32:54