使用lxml解析XML出现意外结果,标签带命名空间前缀
问题原因
你看到的{http://foo.bar}前缀是XML默认命名空间的URI。你的XML里<stationaer>标签带了xmlns="http://foo.bar"声明,这意味着该标签及其所有子元素(除非有其他命名空间声明)都属于这个命名空间。lxml解析时会严格遵循XML命名空间规范,把命名空间URI和标签名组合成{命名空间URI}标签名的格式,所以输出会带这个前缀。
解决方法
如果只想获取不带命名空间的本地标签名,有两种简单方式:
方法1:提取本地名称
lxml的元素对象自带localname属性,直接调用就能拿到纯标签名:
root = etree.fromstring(xml_src) print(repr(root.localname)) # 输出 'stationaer' child = root.getchildren()[0] print(repr(child.localname)) # 输出 'einrichtung'
也可以手动拆分tag字符串(但不如localname规范):
print(repr(root.tag.split('}')[-1]))
方法2:用命名空间映射查询元素
如果后续需要查找指定命名空间下的元素,可以定义命名空间映射,用带前缀的方式精准查询:
ns_map = {'ns': 'http://foo.bar'} # 查找所有einrichtung元素 einrichtungen = root.findall('ns:einrichtung', namespaces=ns_map) for elem in einrichtungen: print(elem.localname) # 输出 'einrichtung'
修改后的完整示例代码
#!/usr/bin/env python3 from lxml import etree xml_src = '''<?xml version="1.0"?> <stationaer xsi:schemaLocation="http:/foo.bar" xmlns="http://foo.bar" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"> <einrichtung> <name>Name</name> </einrichtung> <einrichtung> <name>Name</name> </einrichtung> </stationaer> ''' root = etree.fromstring(xml_src) print(repr(root.localname)) print(repr(root.text)) child = root.getchildren()[0] print(repr(child.localname)) print(repr(child.text))
运行后输出:
'stationaer' ' ' 'einrichtung' ' '
内容的提问来源于stack exchange,提问作者buhtz
相关产品推荐
相关产品推荐

