Python Pandas:如何通过经纬度查找最近点?
问题描述
我是Pandas新手,现在遇到了基于两个DataFrame列数据做关联的问题。我有两个DataFrame:一个是记录道路段的Road Shape DataFrame,另一个是道路附近点位的Points DataFrame。
Road Shape DataFrame
shape_pt_lat shape_pt_lon shape_pt_sequence 2583910 53.402329 -6.150988 1 2583911 53.402334 -6.151043 2 2583912 53.402345 -6.151175 3 2583913 53.402359 -6.151328 4 2583914 53.402518 -6.152953 5 ... ... ... ...
Points DataFrame
latitude longitude timestamp 0 53.376873 -6.216212 1.686826e+09 1 53.370968 -6.223517 1.686827e+09 2 53.363358 -6.234719 1.686827e+09 3 53.360840 -6.238742 1.686827e+09 4 53.355160 -6.246171 1.686827e+09 .. ... ... ...
我想给Points DataFrame新增一列,记录Road Shape DataFrame中距离当前点最近的那一行的索引。如果遍历Points的每个点,再逐个遍历Road的点计算距离找最小值,效率太低了。有没有更优的实现方式?
补充尝试
我试过下面的方案,虽然能跑,但耗时是单纯遍历的4倍左右,有没有办法进一步提速?
shape_tup = [tuple(r) for r in shape_df[['shape_pt_lat', 'shape_pt_lon']].to_numpy()] pos_points["shape_pt_ind"] = np.nan for index, pos in pos_points.iterrows(): min_pair = min(shape_tup, key=lambda t: (abs(t[0] - pos["latitude"]) + abs(t[1] - pos["longitude"]))) min_index = shape_df.index[(shape_df[['shape_pt_lat', 'shape_pt_lon']].values[:, None] == min_pair).all(2).any(1)] pos_points.loc[index, "shape_pt_ind"] = min_index[0]
高效解决方案
用scipy.spatial.KDTree做最近邻搜索是这类问题的最优方案,能把时间复杂度从O(n*m)降到O(n log m),速度提升非常明显。
实现步骤
- 从Road Shape DataFrame提取经纬度数据,构建KDTree;
- 用Points DataFrame的经纬度数据查询KDTree,得到最近邻的位置索引;
- 将位置索引映射回Road的原行索引,添加到Points DataFrame中。
代码示例
import numpy as np from scipy.spatial import KDTree import pandas as pd # 提取道路点的经纬度数组 road_coords = shape_df[['shape_pt_lat', 'shape_pt_lon']].values # 构建KDTree空间索引 kdtree = KDTree(road_coords) # 提取点位的经纬度数组 point_coords = pos_points[['latitude', 'longitude']].values # 查询每个点的最近邻,返回距离和road_coords中的行位置 distances, nearest_pos_indices = kdtree.query(point_coords, k=1) # 将road_coords的行位置映射为shape_df的原索引,添加到Points DataFrame pos_points['shape_pt_ind'] = shape_df.index[nearest_pos_indices]
为什么比你的尝试更快?
你之前的方案有两个核心低效点:
- 用
min遍历所有道路点找最近邻,本质还是暴力遍历的O(n*m)复杂度; - 找到最近点后用布尔索引匹配原DataFrame,多了一次不必要的全量遍历。
KDTree基于空间划分结构,能快速缩小搜索范围,查询效率远高于暴力遍历;同时直接通过索引映射一步到位,省去了后续的匹配步骤。
额外优化(针对超大数据量)
如果数据量达到百万级以上,可以考虑:
- 用
geopandas的空间索引,专门处理地理空间数据的最近邻查询; - 在精度允许的前提下,对道路点做降采样,减少KDTree的节点数量。
内容的提问来源于stack exchange,提问作者M B
相关产品推荐
相关产品推荐

