lxml XPath返回空列表求助:带命名空间XML解析问题
问题
尝试解析字节格式的XML并读取内容,但所有XPath查询均返回空列表。
待解析XML
<?xml version="1.0" encoding="UTF-8"?> <Invoice xmlns="urn:eslog:2.00" xmlns:in="http://uri.etsi.org/01903/v1.1.1#" xmlns:io="http://www.w3.org/2000/09/xmldsig#" xmlns:xs4xs="http://www.w3.org/2001/XMLSchema" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="urn:eslog:2.00 eSLOG20_INVOIC_v200.xsd"> <M_INVOIC Id="data"> <S_UNH> <D_0062>1889</D_0062> <C_S009> <D_0065>INVOIC</D_0065> <D_0052>D</D_0052> <D_0054>01B</D_0054> <D_0051>UN</D_0051> </C_S009> </S_UNH> <S_BGM> <C_C002> <D_1001>380</D_1001> </C_C002> <C_C106> <D_1004>1889</D_1004> </C_C106> </S_BGM> </M_INVOIC> </Invoice>
尝试的代码(注释为输出结果)
def load(self, file): # print(file) # b'<?xml version="1.0" ... # print(type(file)) # <class 'bytes'> root = etree.parse(BytesIO(file)) # print(root.tag) # Returns: object has no attribute 'tag' print(root.xpath('/Invoice')) # Returns: [] # print(root.nsmap) # object has no attribute 'nsmap' root = etree.fromstring(file) print(root.tag) # {urn:eslog:2.00}Invoice # print(root.xpath('/Invoice')) # [] # print(root.xpath('/{urn:eslog:2.00}Invoice')) # Invalid expression print(root.nsmap) # {None: 'urn:eslog:2.00', 'in': ... print(root.xpath('/Invoce/M_INVOIC', nsmap = root.nsmap[None])) # [] for all dict keys ns = { # a copy of root.nsmap 'None': 'urn:eslog:2.00', 'in': 'http://uri.etsi.org/01903/v1.1.1#', 'io': 'http://www.w3.org/2000/09/xmldsig#', 'xs4xs': 'http://www.w3.org/2001/XMLSchema', 'xsi': 'http://www.w3.org/2001/XMLSchema-instance'} print(root.xpath('/Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty [] print(root.xpath('/in:Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty [] print(root.xpath('/io:Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty [] print(root.xpath('/xs4xs:Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty [] print(root.xpath('/xsi:Invoice/M_INVOIC/S_BGM/C_C106/D_1004', namespaces=ns)) # empty []
使用fromstring()能正常返回结果,但所有XPath查询均返回空列表,需解决该问题。
解决方案
问题核心是XML命名空间处理错误,以下是具体修正步骤:
1. 错误原因
- XML根元素
Invoice及所有子元素都属于默认命名空间urn:eslog:2.00,XPath默认只会匹配无命名空间的元素,直接写/Invoice无法命中目标。 - 命名空间字典用
'None'作为键无效,lxml要求默认命名空间必须分配自定义前缀(如'es')。
2. 修正后的代码
def load(self, file): root = etree.fromstring(file) # 给默认命名空间分配自定义前缀,其他命名空间保留原前缀 ns = { 'es': 'urn:eslog:2.00', 'in': 'http://uri.etsi.org/01903/v1.1.1#', 'io': 'http://www.w3.org/2000/09/xmldsig#', 'xs4xs': 'http://www.w3.org/2001/XMLSchema', 'xsi': 'http://www.w3.org/2001/XMLSchema-instance' } # 查询时给所有默认命名空间的元素加上前缀 target_node = root.xpath('/es:Invoice/es:M_INVOIC/es:S_BGM/es:C_C106/es:D_1004', namespaces=ns) if target_node: print(target_node[0].text) # 输出:1889 # 如果用etree.parse()处理,需注意返回的是文档对象,查询逻辑一致 doc = etree.parse(BytesIO(file)) doc_target = doc.xpath('/es:Invoice/es:M_INVOIC/es:S_BGM/es:C_C106/es:D_1004', namespaces=ns) if doc_target: print(doc_target[0].text) # 输出:1889
3. 关键注意事项
- 所有继承默认命名空间的元素,XPath查询时必须添加自定义前缀。
etree.parse()返回的是文档对象,不是根节点,但XPath查询规则和根节点一致。- 不要直接在XPath中写命名空间URI(如
/{urn:eslog:2.00}Invoice),lxml不支持这种语法。
内容的提问来源于stack exchange,提问作者Kaspero
相关产品推荐
相关产品推荐

