Python ElementTree解析XML 按ID提取匹配节点属性值
Python提取OSM XML中way元素数据实现方案
元组格式输出实现代码
基于你已经使用的ElementTree库,只需要在遍历单个way元素时同步收集nd节点ref、匹配maxspeed标签即可,完整可运行代码如下:
import xml.etree.ElementTree as ET # 若需要缺失值用NaN表示,导入math模块即可 import math # 替换为你的XML文件实际路径 tree = ET.parse("your_osm_data.xml") root = tree.getroot() result_list = [] for way in root.findall("way"): # 提取way唯一ID,转为整数类型 way_id = int(way.get("id")) # 收集当前way下所有nd节点的ref值,转为整数后组装为元组 node_ids = tuple(int(nd.get("ref")) for nd in way.findall("nd")) # 查找maxspeed标签值 maxspeed = None # 若要返回NaN替换为 math.nan for tag in way.findall("tag"): if tag.get("k") == "maxspeed": maxspeed = tag.get("v") break # 找到目标标签后终止循环,减少无效遍历 result_list.append((way_id, node_ids, maxspeed)) # 按你要求的格式打印输出 print("(id, node_ids, maxspeed)") for item in result_list: print(item)
基于你提供的示例XML,运行后输出结果如下:
(id, node_ids, maxspeed) (4260867, (25550395, 25550396), 'none') (312407268, (7792523927, 25393142, 5583629192, 25393143), '60') (106141287, (913936737, 1222080363), None)
注:你给出的参考输出中id=4260867的way实际存在
k="maxspeed" v="none"的标签,不属于缺失值,代码会正确提取到字符串'none';只有完全不带maxspeed标签的way才会返回预设的缺失值(None/NaN)。
比元组更合理的数据组织方式
元组需要靠位置索引取值,字段多的时候容易记混顺序,更推荐以下三种方式:
- 字典列表:每个way对应一个键名明确的字典,取数时直接按键名访问,不需要记位置,适合轻量数据处理场景,格式示例:
{"way_id": 4260867, "node_ids": (25550395, 25550396), "maxspeed": "none"} - dataclass结构化对象(Python 3.7+原生支持):可以定义带类型提示的数据类,IDE自动补全更友好,适合后续要写复杂业务逻辑的场景:
from dataclasses import dataclass from typing import Tuple, Optional @dataclass class RoadWay: way_id: int node_ids: Tuple[int, ...] maxspeed: Optional[str]
- Pandas DataFrame:如果后续要做数据统计、筛选、聚合(比如统计不同等级道路的限速分布、导出CSV),直接转为DataFrame效率最高:
import pandas as pd way_df = pd.DataFrame(result_list, columns=["way_id", "node_ids", "maxspeed"]) # 筛选限速60的道路直接用 way_df[way_df["maxspeed"] == "60"] 即可
内容的提问来源于stack exchange,提问作者Kpol
相关产品推荐
相关产品推荐

