如何用Python从带命名空间的XML提取ProductDescription至列表
问题原因及解决方案
你的代码无法找到目标节点的核心原因是XML使用了默认命名空间,根节点的xmlns="urn:OECD:StandardAuditFile-Tax:PT_1.04_01"声明了所有子节点都属于该命名空间,直接用标签名(如Invoice)查找会匹配不到任何节点,因此输出None。
下面是两种可行的解决方案:
方法一:使用命名空间映射(推荐)
通过定义命名空间前缀,精准匹配目标节点:
import xml.etree.ElementTree as ET # 解析XML文件 tree = ET.parse("2023_10_2023.xml") root = tree.getroot() # 定义命名空间映射,前缀可自定义 ns = {'oecd': 'urn:OECD:StandardAuditFile-Tax:PT_1.04_01'} product_descriptions = [] # 按层级查找指定路径的节点 for invoice in root.findall('.//oecd:Invoice', ns): for line in invoice.findall('oecd:Line', ns): pd_element = line.find('oecd:ProductDescription', ns) # 确保节点和文本不为空,避免异常 if pd_element and pd_element.text: product_descriptions.append(pd_element.text) # 输出结果列表 print(product_descriptions)
方法二:忽略命名空间(快速测试用)
通过local-name()匹配标签名,绕过命名空间限制(不推荐用于复杂XML,可能误匹配其他命名空间的节点):
import xml.etree.ElementTree as ET tree = ET.parse("2023_10_2023.xml") root = tree.getroot() product_descriptions = [] # 用local-name()匹配标签名,忽略命名空间 for invoice in root.findall('.//*[local-name()="Invoice"]'): for line in invoice.findall('.//*[local-name()="Line"]'): pd_element = line.find('.//*[local-name()="ProductDescription"]') if pd_element and pd_element.text: product_descriptions.append(pd_element.text) print(product_descriptions)
两种方法最终都会将所有AuditFile/SourceDocuments/SalesInvoices/Invoice/Line/ProductDescription的文本提取到product_descriptions列表中。
内容的提问来源于stack exchange,提问作者Simão Horta
相关产品推荐
相关产品推荐

