经纬度计算2km半径内重叠场馆数量的Python实现方案咨询
实现方案
以下是两种可直接落地的实现方式,你可以根据依赖安装情况选择:
方案1:BallTree 高效实现(推荐)
针对地理空间距离计算场景,scikit-learn的BallTree支持直接使用球面距离(haversine)做近邻查询,1500条数据的计算耗时可以忽略,结果精准:
import pandas as pd import numpy as np from sklearn.neighbors import BallTree # 替换为你自己的数据集读取逻辑 df = pd.read_csv('你的场馆数据集路径.csv') # 先清理经纬度空值 df = df.dropna(subset=['latitude', 'longitude']).reset_index(drop=True) # BallTree要求输入为弧度,顺序为[纬度, 经度] coords = np.radians(df[['latitude', 'longitude']].values) # 构建BallTree,使用haversine距离(返回结果单位为地球半径) tree = BallTree(coords, metric='haversine') # 地球平均半径为6371km,2km对应的弧度值 = 距离/地球半径 radius = 2 / 6371 # 查询每个点半径范围内的所有点数量 counts = tree.query_radius(coords, r=radius, count_only=True) # 减1是排除场馆自身,得到其他场馆的数量 df['overlap'] = counts - 1
方案2:haversine包暴力实现
如果不想额外引入scikit-learn依赖,用haversine包直接遍历计算也完全可行,1500条数据的计算量极低不会有性能问题:
import pandas as pd from haversine import haversine, Unit df = pd.read_csv('你的场馆数据集路径.csv') df = df.dropna(subset=['latitude', 'longitude']).reset_index(drop=True) overlap_counts = [] for i, row in df.iterrows(): current_loc = (row['latitude'], row['longitude']) cnt = 0 for j, other_row in df.iterrows(): # 跳过当前场馆自身 if i == j: continue other_loc = (other_row['latitude'], other_row['longitude']) # 计算距离单位为km dist = haversine(current_loc, other_loc, unit=Unit.KILOMETERS) if dist <= 2: cnt +=1 overlap_counts.append(cnt) df['overlap'] = overlap_counts
注意事项
- 两种方案计算结果完全一致,BallTree的性能优势在数据量超过1万条后会更明显
- 计算时注意统一单位,避免出现米和公里混用的误差
- 如果你的经纬度字段名是
lat/lng,替换代码中的对应字段名即可
内容的提问来源于stack exchange,提问作者Lee Vincent
相关产品推荐
相关产品推荐

