You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算组内经纬度点间距离及统计指定范围内点数

解决方案

首先,你需要在groupby.apply的自定义函数里,先提取组内第一个点的经纬度,再用h3.point_dist逐个计算组内每个点和它的距离,最后统计符合阈值的数量。

完整代码示例

假设你的DataFrame包含col1、col2、latitude、longitude列,以下是可直接运行的实现代码:

import pandas as pd
import h3

# 模拟测试数据
df = pd.DataFrame({
    'col1': [1,2,3,4,1,2,3],
    'col2': [1,1,1,1,2,2,2],
    'latitude': [30.1234, 30.1235, 30.1236, 30.1240, 40.5678, 40.5679, 40.5685],
    'longitude': [120.4567, 120.4568, 120.4569, 120.4575, 110.1234, 110.1235, 110.1240]
})

# 定义分组处理函数
def count_points_within_threshold(group, threshold=2):
    # 取组内第一个点的经纬度(已按col1排序,iloc[0]就是目标第1行)
    ref_lat, ref_lon = group.iloc[0]['latitude'], group.iloc[0]['longitude']
    # 计算组内每个点与参考点的距离(h3.point_dist默认单位为米)
    group['distance'] = group.apply(
        lambda row: h3.point_dist((ref_lat, ref_lon), (row['latitude'], row['longitude'])),
        axis=1
    )
    # 统计距离符合阈值的点数,返回与组行数一致的Series确保对齐
    count = (group['distance'] <= threshold).sum()
    return pd.Series([count]*len(group), index=group.index, name='point_within_range')

# 执行排序、分组、计算逻辑
df = df.sort_values(by=['col1','col2'])
df['point_within_range'] = df.groupby('col2').apply(count_points_within_threshold).reset_index(level=0, drop=True)

print(df)

关键步骤说明

  • 排序与分组:先按col1、col2排序保证组内行顺序符合要求,再按col2完成分组。
  • 提取参考点:通过group.iloc[0]获取组内第一行的经纬度作为基准点。
  • 距离计算:用group.apply遍历组内每行,调用h3.point_dist计算当前点与参考点的距离(注意参数是(纬度, 经度)的元组,顺序不要搞反)。
  • 统计结果:用布尔索引统计距离≤阈值的点数,返回与组行数相同的Series,确保分组计算结果能和原DataFrame的行一一对应。

高效替代方案(适用于大数据量)

如果你的数据行数较多,逐行apply效率偏低,可以用批量方式优化:

def count_points_within_threshold(group, threshold=2):
    ref_lat, ref_lon = group.iloc[0]['latitude'], group.iloc[0]['longitude']
    # 把组内经纬度转为numpy数组批量处理
    coords = group[['latitude', 'longitude']].to_numpy()
    # 批量计算所有点到参考点的距离
    distances = [h3.point_dist((ref_lat, ref_lon), (lat, lon)) for lat, lon in coords]
    count = sum(d <= threshold for d in distances)
    return pd.Series([count]*len(group), index=group.index, name='point_within_range')

这个方法避免了逐行遍历的开销,处理大数据时速度会明显提升。

内容的提问来源于stack exchange,提问作者Prathamesh Sawant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 21:35:20