如何提取m:properties标签之间的数据值?技术实现方法咨询
嘿,要提取XML里<m:properties>标签之间的数据其实挺简单的,根据你使用的工具或编程语言,有几种实用的方法,我给你整理了几个常用的方案:
方法1:用Python标准库
xml.etree.ElementTree 这是Python自带的XML解析工具,不需要额外安装依赖。核心是处理好XML的命名空间(因为你的XML里有m:、d:这类前缀,对应特定的命名空间URI),具体步骤如下:
import xml.etree.ElementTree as ET # 假设你的XML内容已经存在这个变量里(或者可以从文件读取) xml_content = '''<entry xmlns:d="http://schemas.microsoft.com/ado/2007/08/dataservices" xmlns:m="http://schemas.microsoft.com/ado/2007/08/dataservices/metadata"> <id>http://data.treasury.gov/Feed.svc/DailyTreasuryYieldCurveRateData(7086)</id> <title type="text"></title> <updated>2018-04-25T02:39:22Z</updated> <author> <name /> </author> <link rel="edit" title="DailyTreasuryYieldCurveRateDatum" href="DailyTreasuryYieldCurveRateData(7086)" /> <category term="TreasuryDataWarehouseModel.DailyTreasuryYieldCurveRateDatum" scheme="http://schemas.microsoft.com/ado/2007/08/dataservices/scheme" /> <content type="application/xml"> <m:properties> <d:Id m:type="Edm.Int32">7086</d:Id> <d:NEW_DATE m:type="Edm.DateTime">2018-04-24T00:00:00</d:NEW_DATE> <d:BC_1MONTH m:type="Edm.Double">2.05</d:BC_1MONTH> <d:BC_3MONTH m:type="Edm.Double">2.14</d:BC_3MONTH> </m:properties> </content> </entry>''' # 注册命名空间,避免解析时丢失前缀 ET.register_namespace('m', 'http://schemas.microsoft.com/ado/2007/08/dataservices/metadata') ET.register_namespace('d', 'http://schemas.microsoft.com/ado/2007/08/dataservices') # 解析XML字符串 root = ET.fromstring(xml_content) # 定位到<m:properties>元素,注意要传入命名空间映射 properties_elem = root.find('.//m:properties', namespaces={ 'm': 'http://schemas.microsoft.com/ado/2007/08/dataservices/metadata' }) # 提取所有子元素的键值对 extracted_data = {} for child in properties_elem: # 去掉标签里的命名空间前缀,只保留字段名 field_name = child.tag.split('}')[-1] extracted_data[field_name] = child.text print(extracted_data)
运行后会输出一个字典,包含<m:properties>里所有字段和对应的值。
方法2:用Python的
lxml库(更灵活) 如果你需要更强大的XML处理能力(比如复杂的XPath查询),可以用lxml库,它的XPath支持更完善:
首先安装依赖:
pip install lxml
然后编写代码:
from lxml import etree xml_content = '''[同上的XML内容]''' # 解析XML root = etree.fromstring(xml_content) # 定义命名空间映射 ns_map = { 'm': 'http://schemas.microsoft.com/ado/2007/08/dataservices/metadata', 'd': 'http://schemas.microsoft.com/ado/2007/08/dataservices' } # 用XPath直接定位<m:properties>,并提取所有d:前缀的子元素 properties_data = {} for elem in root.xpath('//m:properties/d:*', namespaces=ns_map): field_name = elem.tag.split('}')[-1] properties_data[field_name] = elem.text print(properties_data)
方法3:用命令行工具
xmllint(快速验证) 如果你不想写代码,用系统自带的xmllint(一般Linux/macOS默认有,Windows需要安装libxml2工具)可以快速提取数据:
# 提取<m:properties>下所有子元素 xmllint --xpath '//*[local-name()="properties" and namespace-uri()="http://schemas.microsoft.com/ado/2007/08/dataservices/metadata"]/*' your_xml_file.xml
这里用了local-name()和namespace-uri()来避开前缀的问题,确保能正确匹配到目标元素。
内容的提问来源于stack exchange,提问作者yadav
相关产品推荐
相关产品推荐

