You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从带命名空间的XML提取ProductDescription至列表

问题原因及解决方案

你的代码无法找到目标节点的核心原因是XML使用了默认命名空间,根节点的xmlns="urn:OECD:StandardAuditFile-Tax:PT_1.04_01"声明了所有子节点都属于该命名空间,直接用标签名(如Invoice)查找会匹配不到任何节点,因此输出None。

下面是两种可行的解决方案:

方法一:使用命名空间映射(推荐)

通过定义命名空间前缀,精准匹配目标节点:

import xml.etree.ElementTree as ET

# 解析XML文件
tree = ET.parse("2023_10_2023.xml")
root = tree.getroot()

# 定义命名空间映射,前缀可自定义
ns = {'oecd': 'urn:OECD:StandardAuditFile-Tax:PT_1.04_01'}
product_descriptions = []

# 按层级查找指定路径的节点
for invoice in root.findall('.//oecd:Invoice', ns):
    for line in invoice.findall('oecd:Line', ns):
        pd_element = line.find('oecd:ProductDescription', ns)
        # 确保节点和文本不为空,避免异常
        if pd_element and pd_element.text:
            product_descriptions.append(pd_element.text)

# 输出结果列表
print(product_descriptions)

方法二:忽略命名空间(快速测试用)

通过local-name()匹配标签名,绕过命名空间限制(不推荐用于复杂XML,可能误匹配其他命名空间的节点):

import xml.etree.ElementTree as ET

tree = ET.parse("2023_10_2023.xml")
root = tree.getroot()

product_descriptions = []

# 用local-name()匹配标签名,忽略命名空间
for invoice in root.findall('.//*[local-name()="Invoice"]'):
    for line in invoice.findall('.//*[local-name()="Line"]'):
        pd_element = line.find('.//*[local-name()="ProductDescription"]')
        if pd_element and pd_element.text:
            product_descriptions.append(pd_element.text)

print(product_descriptions)

两种方法最终都会将所有AuditFile/SourceDocuments/SalesInvoices/Invoice/Line/ProductDescription的文本提取到product_descriptions列表中。

内容的提问来源于stack exchange,提问作者Simão Horta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 12:05:03