You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

lxml XPath返回空列表求助:带命名空间XML解析问题

问题

尝试解析字节格式的XML并读取内容,但所有XPath查询均返回空列表。

待解析XML

<?xml version="1.0" encoding="UTF-8"?>
<Invoice
    xmlns="urn:eslog:2.00"
    xmlns:in="http://uri.etsi.org/01903/v1.1.1#"
    xmlns:io="http://www.w3.org/2000/09/xmldsig#"
    xmlns:xs4xs="http://www.w3.org/2001/XMLSchema"
    xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="urn:eslog:2.00 eSLOG20_INVOIC_v200.xsd">
    <M_INVOIC Id="data">
        <S_UNH>
            <D_0062>1889</D_0062>
            <C_S009>
                <D_0065>INVOIC</D_0065>
                <D_0052>D</D_0052>
                <D_0054>01B</D_0054>
                <D_0051>UN</D_0051>
            </C_S009>
        </S_UNH>
        <S_BGM>
            <C_C002>
                <D_1001>380</D_1001>
            </C_C002>
            <C_C106>
                <D_1004>1889</D_1004>
            </C_C106>
        </S_BGM>
    </M_INVOIC>
</Invoice>

尝试的代码(注释为输出结果)

def load(self, file):

        # print(file)  # b'<?xml version="1.0" ...
        # print(type(file))  # <class 'bytes'>

        root = etree.parse(BytesIO(file))
        # print(root.tag)  # Returns: object has no attribute 'tag'
        print(root.xpath('/Invoice'))  # Returns: []
        # print(root.nsmap)  # object has no attribute 'nsmap'

        root = etree.fromstring(file)
        print(root.tag)  # {urn:eslog:2.00}Invoice
        # print(root.xpath('/Invoice'))  # []
        # print(root.xpath('/{urn:eslog:2.00}Invoice')) # Invalid expression
        print(root.nsmap)  # {None: 'urn:eslog:2.00', 'in': ...
        print(root.xpath('/Invoce/M_INVOIC', nsmap = root.nsmap[None]))  # [] for all dict keys
        ns = { # a copy of root.nsmap
            'None': 'urn:eslog:2.00',
            'in': 'http://uri.etsi.org/01903/v1.1.1#',
            'io': 'http://www.w3.org/2000/09/xmldsig#',
            'xs4xs': 'http://www.w3.org/2001/XMLSchema',
            'xsi': 'http://www.w3.org/2001/XMLSchema-instance'}
        print(root.xpath('/Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty []
        print(root.xpath('/in:Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty []
        print(root.xpath('/io:Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty []
        print(root.xpath('/xs4xs:Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty []
        print(root.xpath('/xsi:Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty []

使用fromstring()能正常返回结果,但所有XPath查询均返回空列表,需解决该问题。


解决方案

问题核心是XML命名空间处理错误,以下是具体修正步骤:

1. 错误原因

  • XML根元素Invoice及所有子元素都属于默认命名空间urn:eslog:2.00,XPath默认只会匹配无命名空间的元素,直接写/Invoice无法命中目标。
  • 命名空间字典用'None'作为键无效,lxml要求默认命名空间必须分配自定义前缀(如'es')。

2. 修正后的代码

def load(self, file):
    root = etree.fromstring(file)
    
    # 给默认命名空间分配自定义前缀,其他命名空间保留原前缀
    ns = {
        'es': 'urn:eslog:2.00',
        'in': 'http://uri.etsi.org/01903/v1.1.1#',
        'io': 'http://www.w3.org/2000/09/xmldsig#',
        'xs4xs': 'http://www.w3.org/2001/XMLSchema',
        'xsi': 'http://www.w3.org/2001/XMLSchema-instance'
    }
    
    # 查询时给所有默认命名空间的元素加上前缀
    target_node = root.xpath('/es:Invoice/es:M_INVOIC/es:S_BGM/es:C_C106/es:D_1004', namespaces=ns)
    if target_node:
        print(target_node[0].text)  # 输出:1889
    
    # 如果用etree.parse()处理,需注意返回的是文档对象,查询逻辑一致
    doc = etree.parse(BytesIO(file))
    doc_target = doc.xpath('/es:Invoice/es:M_INVOIC/es:S_BGM/es:C_C106/es:D_1004', namespaces=ns)
    if doc_target:
        print(doc_target[0].text)  # 输出:1889

3. 关键注意事项

  • 所有继承默认命名空间的元素,XPath查询时必须添加自定义前缀。
  • etree.parse()返回的是文档对象,不是根节点,但XPath查询规则和根节点一致。
  • 不要直接在XPath中写命名空间URI(如/{urn:eslog:2.00}Invoice),lxml不支持这种语法。

内容的提问来源于stack exchange,提问作者Kaspero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 16:01:07