如何快速为DataFrame中每个点获取指定半径内的点列表?
快速查找经纬度点指定半径内的邻接点列表
针对你的需求,最快的实现方式是利用空间索引避免全量距离计算,下面提供两种高效方案:
方法一:使用GeoPandas(直观易用)
GeoPandas专门处理空间数据,自带空间索引,能大幅提升邻点查询效率。
步骤与代码:
- 导入依赖库并创建DataFrame
import pandas as pd import geopandas as gpd from shapely.geometry import Point # 初始化你的数据 df = pd.DataFrame({ 'projectName': ['a', 'b', 'c', 'd', 'e'], 'latitude': [56.864229, 55.810413, 55.924912, 56.804987, 55.806000], 'longitude': [60.609576, 37.701168, 37.966033, 60.590667, 37.569863] })
- 转换为GeoDataFrame并构建空间索引
# 将经纬度转为地理点(注意GeoPandas要求先经度后纬度) gdf = gpd.GeoDataFrame( df, geometry=gpd.points_from_xy(df.longitude, df.latitude), crs='EPSG:4326' # 使用WGS84地理坐标系 ) # 构建空间索引,用于快速候选点筛选 sindex = gdf.sindex
- 定义邻点查询函数并生成结果
radius = 30000 # 30公里转为米(GeoPandas距离计算默认单位为米) def get_neighbors(row): # 先用空间索引筛选边界框内的候选点(比全量计算快数倍) candidate_ids = list(sindex.intersection(row.geometry.buffer(radius).bounds)) candidates = gdf.iloc[candidate_ids] # 精确计算球面距离并筛选符合条件的点(排除自身) valid_neighbors = candidates[ (candidates.distance(row.geometry) <= radius) & (candidates.projectName != row.projectName) ] return valid_neighbors['projectName'].tolist() # 生成30km邻点列 gdf['30km'] = gdf.apply(get_neighbors, axis=1) # 输出结果 print(gdf[['projectName', 'latitude', 'longitude', '30km']])
方法二:使用SciPy cKDTree(高效轻量)
如果不想引入GeoPandas依赖,可使用SciPy的cKDTree实现球面距离的高效查询,适合大数据量场景。
步骤与代码:
- 导入依赖库并初始化数据
import pandas as pd import numpy as np from scipy.spatial import cKDTree df = pd.DataFrame({ 'projectName': ['a', 'b', 'c', 'd', 'e'], 'latitude': [56.864229, 55.810413, 55.924912, 56.804987, 55.806000], 'longitude': [60.609576, 37.701168, 37.966033, 60.590667, 37.569863] })
- 转换坐标并构建KD树
# 将经纬度转为弧度,用于球面距离计算 coords_rad = np.radians(df[['latitude', 'longitude']].values) # 转换为球面坐标系(colatitude = 90°-纬度,单位弧度) coords_spherical = np.column_stack((np.pi/2 - coords_rad[:, 0], coords_rad[:, 1])) # 构建KD树 tree = cKDTree(coords_spherical) # 将30公里转为弧度距离(地球半径取6371公里) radius_rad = 30 / 6371.0
- 查询邻点并整理结果
# 获取每个点的邻点索引 neighbor_indices = tree.query_ball_point(coords_spherical, radius_rad) # 转换为projectName列表并排除自身 df['30km'] = [ df.iloc[idx]['projectName'].tolist() for idx in neighbor_indices ] df['30km'] = df.apply(lambda row: [name for name in row['30km'] if name != row['projectName']], axis=1) # 输出结果 print(df)
两种方法都能快速得到你需要的结果,其中GeoPandas更适合后续还有空间数据处理需求的场景,而SciPy方案则更轻量高效。
内容的提问来源于stack exchange,提问作者Kirill Kondratenko
相关产品推荐
相关产品推荐

