You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas:如何通过经纬度查找最近点?

问题描述

我是Pandas新手,现在遇到了基于两个DataFrame列数据做关联的问题。我有两个DataFrame:一个是记录道路段的Road Shape DataFrame,另一个是道路附近点位的Points DataFrame。

Road Shape DataFrame

shape_pt_lat  shape_pt_lon  shape_pt_sequence
2583910     53.402329     -6.150988                  1
2583911     53.402334     -6.151043                  2
2583912     53.402345     -6.151175                  3
2583913     53.402359     -6.151328                  4
2583914     53.402518     -6.152953                  5
...               ...           ...                ...

Points DataFrame

latitude  longitude     timestamp
0   53.376873  -6.216212  1.686826e+09
1   53.370968  -6.223517  1.686827e+09
2   53.363358  -6.234719  1.686827e+09
3   53.360840  -6.238742  1.686827e+09
4   53.355160  -6.246171  1.686827e+09
..        ...        ...           ...

我想给Points DataFrame新增一列,记录Road Shape DataFrame中距离当前点最近的那一行的索引。如果遍历Points的每个点,再逐个遍历Road的点计算距离找最小值,效率太低了。有没有更优的实现方式?

补充尝试

我试过下面的方案,虽然能跑,但耗时是单纯遍历的4倍左右,有没有办法进一步提速?

shape_tup = [tuple(r) for r in shape_df[['shape_pt_lat', 'shape_pt_lon']].to_numpy()]

pos_points["shape_pt_ind"] = np.nan

for index, pos in pos_points.iterrows():
    min_pair = min(shape_tup, key=lambda t: (abs(t[0] - pos["latitude"]) + abs(t[1] - pos["longitude"])))
    min_index = shape_df.index[(shape_df[['shape_pt_lat', 'shape_pt_lon']].values[:, None] == min_pair).all(2).any(1)]
    pos_points.loc[index, "shape_pt_ind"] = min_index[0]

高效解决方案

用scipy.spatial.KDTree做最近邻搜索是这类问题的最优方案,能把时间复杂度从O(n*m)降到O(n log m),速度提升非常明显。

实现步骤

  1. 从Road Shape DataFrame提取经纬度数据,构建KDTree;
  2. 用Points DataFrame的经纬度数据查询KDTree,得到最近邻的位置索引;
  3. 将位置索引映射回Road的原行索引,添加到Points DataFrame中。

代码示例

import numpy as np
from scipy.spatial import KDTree
import pandas as pd

# 提取道路点的经纬度数组
road_coords = shape_df[['shape_pt_lat', 'shape_pt_lon']].values
# 构建KDTree空间索引
kdtree = KDTree(road_coords)

# 提取点位的经纬度数组
point_coords = pos_points[['latitude', 'longitude']].values
# 查询每个点的最近邻,返回距离和road_coords中的行位置
distances, nearest_pos_indices = kdtree.query(point_coords, k=1)

# 将road_coords的行位置映射为shape_df的原索引,添加到Points DataFrame
pos_points['shape_pt_ind'] = shape_df.index[nearest_pos_indices]

为什么比你的尝试更快?

你之前的方案有两个核心低效点:

  • 用min遍历所有道路点找最近邻,本质还是暴力遍历的O(n*m)复杂度;
  • 找到最近点后用布尔索引匹配原DataFrame,多了一次不必要的全量遍历。

KDTree基于空间划分结构,能快速缩小搜索范围,查询效率远高于暴力遍历;同时直接通过索引映射一步到位,省去了后续的匹配步骤。

额外优化(针对超大数据量)

如果数据量达到百万级以上,可以考虑:

  • 用geopandas的空间索引,专门处理地理空间数据的最近邻查询;
  • 在精度允许的前提下,对道路点做降采样,减少KDTree的节点数量。

内容的提问来源于stack exchange,提问作者M B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 17:20:38