如何用SimpleXML获取标签中的属性?附XML结构示例
从指定XML站点提取数据的实用方案
嘿,我来帮你搞定从那个XML站点抓取数据的事儿!先给你捋捋具体的实现思路和代码示例,完全贴合你给出的XML结构~
1. 工具选择
咱们用Python来实现最方便,核心用到两个库:
requests:用来发送HTTP请求获取XML内容xml.etree.ElementTree:Python内置的XML解析库(如果需要更强大的查询能力,也可以用lxml)
2. 具体实现步骤
第一步:获取XML内容
先把XML数据从站点拉下来,代码示例如下:
import requests from xml.etree import ElementTree as ET # 替换成你的目标XML站点URL target_xml_url = "你的XML站点地址" # 发送请求获取XML内容 try: response = requests.get(target_xml_url) response.raise_for_status() # 如果请求失败会抛出异常 except requests.exceptions.RequestException as e: print(f"请求XML站点失败: {e}")
第二步:解析并提取目标数据
根据你给出的XML结构,<body>下包含多个<predictions>节点,有的带<message>,有的带<direction>和<prediction>,咱们针对性解析:
# 解析XML内容 root = ET.fromstring(response.content) # 遍历所有<predictions>节点 for predictions_node in root.findall('predictions'): # 提取message节点的text属性 message_node = predictions_node.find('message') if message_node is not None: print(f"提示消息: {message_node.get('text')}") # 提取direction和prediction的属性 direction_node = predictions_node.find('direction') if direction_node is not None: print(f"方向名称: {direction_node.get('title')}") # 遍历该方向下的所有prediction节点 for prediction_node in direction_node.findall('prediction'): print(f" 车辆编号: {prediction_node.get('vehicle')}") print(f" 剩余时间(分钟): {prediction_node.get('minutes')}") print(f" 时间戳: {prediction_node.get('epochTime')}") # 其他属性比如seconds、isDeparture等都可以用get()方法提取
3. 进阶优化(可选)
如果你的XML结构更复杂,或者需要更灵活的查询,可以用lxml库的XPath功能,示例如下:
from lxml import etree # 用lxml解析XML tree = etree.fromstring(response.content) # 用XPath直接定位所有prediction节点 all_predictions = tree.xpath('//predictions/direction/prediction') for pred in all_predictions: print(f"车辆[{pred.get('vehicle')}] 剩余{pred.get('minutes')}分钟出发")
内容的提问来源于stack exchange,提问作者Hawkydoky
相关产品推荐
相关产品推荐

