You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析带命名空间的XML并提取item节点值

解析带命名空间的XML并提取item节点值的Python方案

目标XML示例

<XMLRESPONSE>
    <SOAP-ENV:Envelope xmlns:SOAP-ENV="http://schemas.xmlsoap.org/soap/envelope/" xmlns:xsd="http://www.w3.org/2001/XMLSchema" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:SOAP-ENC="http://schemas.xmlsoap.org/soap/encoding/" xmlns:tns="urn:wsNotes" SOAP-ENV:encodingStyle="http://schemas.xmlsoap.org/soap/encoding/">
        <SOAP-ENV:Body>
            <ns1:GetPersonasResponse xmlns:ns1="wsNotes">
                <return xsi:type="SOAP-ENC:Array" SOAP-ENC:arrayType="tns:registro[1]">
                    <item xsi:type="tns:registro">
                        <seguro xsi:type="xsd:string">0</seguro>
                        <cedula_pasaporte xsi:type="xsd:string">x-xxx-1454</cedula_pasaporte>
                        <nombre xsi:type="xsd:string">JUANITCO CARDENAS</nombre>
                        <razon_social xsi:type="xsd:string">GOOGLE</razon_social>
                        <patrono xsi:type="xsd:string">GOOGLE PA</patrono>
                        <ruc xsi:nil="true" xsi:type="xsd:string"/>
                        <direccion xsi:type="xsd:string">AVE. MEXICO Y CL. 33 LOCAL. 07</direccion>
                        <telefono1 xsi:type="xsd:int">2259444</telefono1>
                        <telefono2 xsi:nil="true" xsi:type="xsd:int"/>
                        <fecha xsi:type="xsd:string">1220</fecha>
                        <salario xsi:type="xsd:decimal">1000</salario>
                        <promedio_salarial xsi:type="xsd:string">1000</promedio_salarial>
                        <Seis_Meses_Mas xsi:type="xsd:string">Si</Seis_Meses_Mas>
                        <Cantidad_Meses xsi:type="xsd:int">15</Cantidad_Meses>
                        <Historial xsi:type="xsd:string">
                                    Fecha: 1220 Patrono: "APPTIVIDAD" Salario: 5500.00||
                                    Fecha: 1120 Patrono: "APPTIVIDAD" Salario: 5500.00||
                                    Fecha: 0920.0 Patrono: APPTIVIDAD Salario: 5500.00||
                                    Fecha: 0820 Patrono: APPTIVIDAD Salario: 5500.35||
                                    Fecha: 0720 Patrono: APPTIVIDAD Salario: 5500.20||
                                    Fecha: 0620 Patrono: APPTIVIDAD Salario: 5500.01||
                                    Fecha: 0420 Patrono: APPTIVIDAD Salario: 5500.22||
                                    Fecha: 0320 Patrono: APPTIVIDAD Salario: 5500.70||
                                    Fecha: 0120 Patrono: APPTIVIDAD Salario: 5500.97||
                                    Fecha: 1219 Patrono: APPTIVIDAD Salario: 5500.82||
                                    Fecha: 1119 Patrono: APPTIVIDAD Salario: 5500.33||
                                    Fecha: 0919 Patrono: APPTIVIDAD Salario: 5500.25|
                        </Historial>
                        <Total_Empleados xsi:type="xsd:int">20</Total_Empleados>
                    </item>
                </return>
            </ns1:GetPersonasResponse>
        </SOAP-ENV:Body>
    </SOAP-ENV:Envelope>
</XMLRESPONSE>

方法一:使用Python标准库xml.etree.ElementTree

处理带命名空间的XML核心是定义命名空间字典,通过前缀映射对应的URI,再用XPath定位节点。

import xml.etree.ElementTree as ET

# 1. 定义命名空间字典,对应XML中的所有命名空间
namespaces = {
    'soap-env': 'http://schemas.xmlsoap.org/soap/envelope/',
    'xsd': 'http://www.w3.org/2001/XMLSchema',
    'xsi': 'http://www.w3.org/2001/XMLSchema-instance',
    'soap-enc': 'http://schemas.xmlsoap.org/soap/encoding/',
    'tns': 'urn:wsNotes',
    'ns1': 'wsNotes'
}

# 2. 解析XML(如果是文件,替换为ET.parse('your_file.xml'))
xml_content = """上述XML字符串内容,或从文件读取"""
root = ET.fromstring(xml_content)

# 3. 定位item节点,使用XPath并传入命名空间
item_node = root.find('.//item', namespaces)

# 4. 提取item节点内的所有子节点值
item_data = {}
for child in item_node:
    # 获取节点标签名(去掉可能的命名空间前缀)
    tag = child.tag.split('}')[-1] if '}' in child.tag else child.tag
    # 处理xsi:nil="true"的空节点
    if child.get(f'{{{namespaces["xsi"]}}}nil') == 'true':
        item_data[tag] = None
    else:
        # 清理文本中的多余空白(针对Historial这类多行文本)
        item_data[tag] = child.text.strip() if child.text else None

# 打印结果
for key, value in item_data.items():
    print(f"{key}: {value}")

方法二:使用lxml库(更灵活的XML处理)

lxml支持更完整的XPath语法,处理命名空间的方式类似,适合复杂XML场景。

首先安装依赖:

pip install lxml

代码示例:

from lxml import etree

# 定义命名空间字典
namespaces = {
    'soap-env': 'http://schemas.xmlsoap.org/soap/envelope/',
    'xsd': 'http://www.w3.org/2001/XMLSchema',
    'xsi': 'http://www.w3.org/2001/XMLSchema-instance',
    'soap-enc': 'http://schemas.xmlsoap.org/soap/encoding/',
    'tns': 'urn:wsNotes',
    'ns1': 'wsNotes'
}

# 解析XML
xml_content = """上述XML字符串内容,或从文件读取"""
root = etree.fromstring(xml_content)

# 定位item节点(XPath写法更灵活)
item_node = root.xpath('//item', namespaces=namespaces)[0]

# 提取子节点数据
item_data = {}
for child in item_node:
    tag = child.tag.split('}')[-1]
    # 检查xsi:nil属性
    if child.get(f'{{{namespaces["xsi"]}}}nil') == 'true':
        item_data[tag] = None
    else:
        item_data[tag] = child.text.strip() if child.text else None

# 输出结果
for key, value in item_data.items():
    print(f"{key}: {value}")

关键说明

  • 命名空间字典必须与XML中定义的xmlns:*完全对应,否则XPath无法定位节点
  • 处理xsi:nil="true"的节点时,需要判断该属性并设置为None,避免提取空文本
  • 对于多行文本(如Historial),使用strip()清理多余的换行和空格,也可根据需求保留原始格式

内容的提问来源于stack exchange,提问作者Bryam Williams Hirsch Trujillo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 05:30:42