You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的lxml获取XML中的xsi:noNamespaceSchemaLocation

如何用Python的lxml获取XML中xsi:noNamespaceSchemaLocation属性的值?

问题描述

我尝试基于xsi:noNamespaceSchemaLocation验证XML,调研后未找到可行方案。我的XML文件内容来自W3School示例:

<shiporder orderid="889923" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="shiporder.xsd">
  <orderperson>John Smith</orderperson>
  <shipto>
    <name>Ola Nordmann</name>
    <address>Langgt 23</address>
    <city>4000 Stavanger</city>
    <country>Norway</country>
  </shipto>
  <item>
    <title>Empire Burlesque</title>
    <note>Special Edition</note>
    <quantity>1</quantity>
    <price>10.90</price>
  </item>
  <item>
    <title>Hide your heart</title>
    <quantity>1</quantity>
    <price>9.90</price>
  </item>
</shiporder>

解析根节点属性后得到结果:{'{http://www.w3.org/2001/XMLSchema-instance}noNamespaceSchemaLocation': 'shiporder.xsd'},请问如何用Python的lxml正确获取该属性值?


解决方案

在lxml中处理带命名空间的属性,核心是要对应上xsi前缀背后的官方URI:http://www.w3.org/2001/XMLSchema-instance。这里有两种简单的方法可以获取目标属性值:

方法1:直接使用Clark表示法的属性名调用get()

lxml在解析带命名空间的属性时,会自动将前缀展开为完整的URI(也就是你看到的带大括号的键),所以直接用这个完整键名调用get()方法就能拿到值:

from lxml import etree

# 解析XML字符串(如果是文件,改用etree.parse("your_file.xml").getroot())
xml_str = """<shiporder orderid="889923" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="shiporder.xsd">
  <orderperson>John Smith</orderperson>
  <shipto>
    <name>Ola Nordmann</name>
    <address>Langgt 23</address>
    <city>4000 Stavanger</city>
    <country>Norway</country>
  </shipto>
  <item>
    <title>Empire Burlesque</title>
    <note>Special Edition</note>
    <quantity>1</quantity>
    <price>10.90</price>
  </item>
  <item>
    <title>Hide your heart</title>
    <quantity>1</quantity>
    <price>9.90</price>
  </item>
</shiporder>"""
root = etree.fromstring(xml_str)

# 用展开后的完整属性名获取值
schema_location = root.get('{http://www.w3.org/2001/XMLSchema-instance}noNamespaceSchemaLocation')
print(schema_location)  # 输出: shiporder.xsd

方法2:定义命名空间字典,简化属性访问

如果觉得Clark表示法太长,你可以定义一个命名空间字典,把xsi前缀映射到对应的URI,然后用更直观的方式获取属性:

from lxml import etree

xml_str = """<shiporder orderid="889923" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="shiporder.xsd">
  <orderperson>John Smith</orderperson>
  <shipto>
    <name>Ola Nordmann</name>
    <address>Langgt 23</address>
    <city>4000 Stavanger</city>
    <country>Norway</country>
  </shipto>
  <item>
    <title>Empire Burlesque</title>
    <note>Special Edition</note>
    <quantity>1</quantity>
    <price>10.90</price>
  </item>
  <item>
    <title>Hide your heart</title>
    <quantity>1</quantity>
    <price>9.90</price>
  </item>
</shiporder>"""
root = etree.fromstring(xml_str)

# 定义命名空间映射
namespaces = {'xsi': 'http://www.w3.org/2001/XMLSchema-instance'}

# 方式A:结合命名空间字典调用get()
schema_location = root.get('xsi:noNamespaceSchemaLocation', namespaces=namespaces)

# 方式B:用XPath表达式结合命名空间字典查询
schema_location = root.xpath('./@xsi:noNamespaceSchemaLocation', namespaces=namespaces)[0]

print(schema_location)  # 输出: shiporder.xsd

说明

你看到的带大括号的属性键是Clark表示法,这是XML命名空间的标准表示方式,lxml会用这种格式来存储带命名空间的属性。只要对应上正确的URI,不管用哪种方法都能轻松获取到noNamespaceSchemaLocation的值。

内容的提问来源于stack exchange,提问作者Dmytro Chasovskyi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:52:53