如何计算组内经纬度点间距离及统计指定范围内点数
解决方案
首先,你需要在groupby.apply的自定义函数里,先提取组内第一个点的经纬度,再用h3.point_dist逐个计算组内每个点和它的距离,最后统计符合阈值的数量。
完整代码示例
假设你的DataFrame包含col1、col2、latitude、longitude列,以下是可直接运行的实现代码:
import pandas as pd import h3 # 模拟测试数据 df = pd.DataFrame({ 'col1': [1,2,3,4,1,2,3], 'col2': [1,1,1,1,2,2,2], 'latitude': [30.1234, 30.1235, 30.1236, 30.1240, 40.5678, 40.5679, 40.5685], 'longitude': [120.4567, 120.4568, 120.4569, 120.4575, 110.1234, 110.1235, 110.1240] }) # 定义分组处理函数 def count_points_within_threshold(group, threshold=2): # 取组内第一个点的经纬度(已按col1排序,iloc[0]就是目标第1行) ref_lat, ref_lon = group.iloc[0]['latitude'], group.iloc[0]['longitude'] # 计算组内每个点与参考点的距离(h3.point_dist默认单位为米) group['distance'] = group.apply( lambda row: h3.point_dist((ref_lat, ref_lon), (row['latitude'], row['longitude'])), axis=1 ) # 统计距离符合阈值的点数,返回与组行数一致的Series确保对齐 count = (group['distance'] <= threshold).sum() return pd.Series([count]*len(group), index=group.index, name='point_within_range') # 执行排序、分组、计算逻辑 df = df.sort_values(by=['col1','col2']) df['point_within_range'] = df.groupby('col2').apply(count_points_within_threshold).reset_index(level=0, drop=True) print(df)
关键步骤说明
- 排序与分组:先按
col1、col2排序保证组内行顺序符合要求,再按col2完成分组。 - 提取参考点:通过
group.iloc[0]获取组内第一行的经纬度作为基准点。 - 距离计算:用
group.apply遍历组内每行,调用h3.point_dist计算当前点与参考点的距离(注意参数是(纬度, 经度)的元组,顺序不要搞反)。 - 统计结果:用布尔索引统计距离≤阈值的点数,返回与组行数相同的Series,确保分组计算结果能和原DataFrame的行一一对应。
高效替代方案(适用于大数据量)
如果你的数据行数较多,逐行apply效率偏低,可以用批量方式优化:
def count_points_within_threshold(group, threshold=2): ref_lat, ref_lon = group.iloc[0]['latitude'], group.iloc[0]['longitude'] # 把组内经纬度转为numpy数组批量处理 coords = group[['latitude', 'longitude']].to_numpy() # 批量计算所有点到参考点的距离 distances = [h3.point_dist((ref_lat, ref_lon), (lat, lon)) for lat, lon in coords] count = sum(d <= threshold for d in distances) return pd.Series([count]*len(group), index=group.index, name='point_within_range')
这个方法避免了逐行遍历的开销,处理大数据时速度会明显提升。
内容的提问来源于stack exchange,提问作者Prathamesh Sawant
相关产品推荐
相关产品推荐

