如何在HPC集群中使用OSMnx优化路网路径查询脚本执行效率
基于Ubuntu Server HPC集群的Python路径计算脚本优化方案
需求说明
现有用于路网最短路径计算的Python脚本,在非集群服务器上单程执行耗时约4分钟,需适配基于Ubuntu Server的HPC集群运行以缩短执行时间,要求支持离线运行。
脚本功能
接收起止点经纬度作为输入,读取预生成的墨西哥哈利斯科州(面积78588km²)驾车路网graphml文件,计算两点间的最短行驶路径、行驶距离。
原始代码
import sys import networkx as nx import osmnx as ox import geopandas as gpd import math from shapely.geometry import Point import json ox.config(use_cache=False) origlatTmp = float(sys.argv[1]) origlngTmp = float(sys.argv[2]) destlatTmp = float(sys.argv[3]) destlngTmp = float(sys.argv[4]) def fromAtoBpoints(origlat, origlng, destlat, destlng): G = ox.load_graphml('/var/www/html/jalisco.graphml') # 该.graphml文件生成逻辑如下: # import osmnx as ox # ox.config(use_cache=True, log_console=True) # G = ox.graph_from_place('Jalisco,Mexico', network_type = 'drive', simplify=False) # G = ox.add_edge_speeds(G) # G = ox.add_edge_travel_times(G) # ox.save_graphml(G, '/var/www/html/jalisco.graphml') # print("Done!") # 墨西哥哈利斯科州面积为78588 km² lats = [] lngs = [] lats.insert(0, origlat) lats.insert(1, destlat) lngs.insert(0, origlng) lngs.insert(1, destlng) points_list = [Point((lng, lat)) for lat, lng in zip(lats, lngs)] points = gpd.GeoSeries(points_list, crs='epsg:4326') points_proj = points.to_crs(G.graph['crs']) nearest_nodes = [ox.distance.nearest_nodes(G, pt.x, pt.y) for pt in points_proj] route = nx.shortest_path(G, nearest_nodes[0], nearest_nodes[1], weight='length') time = nx.shortest_path_length(G, nearest_nodes[0], nearest_nodes[1], weight='travel_time') # print("Tiempo:",time/60,"min") distance = nx.shortest_path_length(G, nearest_nodes[0], nearest_nodes[1], weight='length') # print("Distancia:",distance/1000,"km") resultArray = [] for a in route: resultArray.append("{lat:"+ str(G.nodes[a]['y'])+",lng:"+ str(G.nodes[a]['x'])+"}") return (distance/1000),"@@@",resultArray print(fromAtoBpoints(origlatTmp,origlngTmp,destlatTmp,destlngTmp))
集群适配优化方案
- 核心冗余逻辑优化:原始脚本最大的性能开销来自每次调用都重复加载graphml路网文件,修改为计算节点启动任务时预加载一次路网到内存,批量处理所有分配的查询请求,单次查询无需重复加载文件,可节省90%以上的执行时间
- 集群调度适配:使用SLURM作为HPC集群调度器,批量查询任务可拆分为多个子任务分配到不同计算节点并行执行,单节点处理的查询数量可根据节点内存配置调整
- 离线运行适配:提前在所有计算节点离线安装
networkx、osmnx、geopandas、shapely等依赖包,将预生成的jalisco.graphml文件同步到所有计算节点本地存储,或者部署到集群共享存储,避免跨节点读取的IO开销 - 算法优化:将默认的Dijkstra最短路径算法替换为双向A*算法,可进一步缩短单次路径查询的耗时;也可预先对哈利斯科州路网做空间分区索引,缩小路径查询的检索范围
- 输出逻辑优化:原始脚本的自定义字符串拼接输出可改为直接输出JSON格式,减少序列化开销,也方便后续结果解析
内容的提问来源于stack exchange,提问作者Juan Martin Gonzalez Razo
相关产品推荐
相关产品推荐

