如何用Python的lxml获取XML中的xsi:noNamespaceSchemaLocation
如何用Python的lxml获取XML中xsi:noNamespaceSchemaLocation属性的值?
问题描述
我尝试基于xsi:noNamespaceSchemaLocation验证XML,调研后未找到可行方案。我的XML文件内容来自W3School示例:
<shiporder orderid="889923" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="shiporder.xsd"> <orderperson>John Smith</orderperson> <shipto> <name>Ola Nordmann</name> <address>Langgt 23</address> <city>4000 Stavanger</city> <country>Norway</country> </shipto> <item> <title>Empire Burlesque</title> <note>Special Edition</note> <quantity>1</quantity> <price>10.90</price> </item> <item> <title>Hide your heart</title> <quantity>1</quantity> <price>9.90</price> </item> </shiporder>
解析根节点属性后得到结果:{'{http://www.w3.org/2001/XMLSchema-instance}noNamespaceSchemaLocation': 'shiporder.xsd'},请问如何用Python的lxml正确获取该属性值?
解决方案
在lxml中处理带命名空间的属性,核心是要对应上xsi前缀背后的官方URI:http://www.w3.org/2001/XMLSchema-instance。这里有两种简单的方法可以获取目标属性值:
方法1:直接使用Clark表示法的属性名调用get()
lxml在解析带命名空间的属性时,会自动将前缀展开为完整的URI(也就是你看到的带大括号的键),所以直接用这个完整键名调用get()方法就能拿到值:
from lxml import etree # 解析XML字符串(如果是文件,改用etree.parse("your_file.xml").getroot()) xml_str = """<shiporder orderid="889923" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="shiporder.xsd"> <orderperson>John Smith</orderperson> <shipto> <name>Ola Nordmann</name> <address>Langgt 23</address> <city>4000 Stavanger</city> <country>Norway</country> </shipto> <item> <title>Empire Burlesque</title> <note>Special Edition</note> <quantity>1</quantity> <price>10.90</price> </item> <item> <title>Hide your heart</title> <quantity>1</quantity> <price>9.90</price> </item> </shiporder>""" root = etree.fromstring(xml_str) # 用展开后的完整属性名获取值 schema_location = root.get('{http://www.w3.org/2001/XMLSchema-instance}noNamespaceSchemaLocation') print(schema_location) # 输出: shiporder.xsd
方法2:定义命名空间字典,简化属性访问
如果觉得Clark表示法太长,你可以定义一个命名空间字典,把xsi前缀映射到对应的URI,然后用更直观的方式获取属性:
from lxml import etree xml_str = """<shiporder orderid="889923" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="shiporder.xsd"> <orderperson>John Smith</orderperson> <shipto> <name>Ola Nordmann</name> <address>Langgt 23</address> <city>4000 Stavanger</city> <country>Norway</country> </shipto> <item> <title>Empire Burlesque</title> <note>Special Edition</note> <quantity>1</quantity> <price>10.90</price> </item> <item> <title>Hide your heart</title> <quantity>1</quantity> <price>9.90</price> </item> </shiporder>""" root = etree.fromstring(xml_str) # 定义命名空间映射 namespaces = {'xsi': 'http://www.w3.org/2001/XMLSchema-instance'} # 方式A:结合命名空间字典调用get() schema_location = root.get('xsi:noNamespaceSchemaLocation', namespaces=namespaces) # 方式B:用XPath表达式结合命名空间字典查询 schema_location = root.xpath('./@xsi:noNamespaceSchemaLocation', namespaces=namespaces)[0] print(schema_location) # 输出: shiporder.xsd
说明
你看到的带大括号的属性键是Clark表示法,这是XML命名空间的标准表示方式,lxml会用这种格式来存储带命名空间的属性。只要对应上正确的URI,不管用哪种方法都能轻松获取到noNamespaceSchemaLocation的值。
内容的提问来源于stack exchange,提问作者Dmytro Chasovskyi
相关产品推荐
相关产品推荐

