You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按tripId分组计算相邻GPS坐标距离并新增Distance列到DataFrame

实现方案

完全可以新增Distance列存储计算结果,你原有代码没有得到预期结果主要有两个问题:

  1. 循环中直接写df.distance=haversine(...)会给整个distance列的所有行覆盖为当前计算值,不是仅给第i行赋值
  2. 没有做行程边界隔离,不同tripId的首行和上一个行程的末行不会计算距离,也不会出现越界问题

正确实现代码

import pandas as pd
import math

def haversine(coord1, coord2):
    R = 6372800 # 地球半径,单位为米
    # 无下一个坐标时返回空值
    if pd.isna(coord2[0]) or pd.isna(coord2[1]):
        return float('nan')
    lat1, lon1 = coord1
    lat2, lon2 = coord2
    
    phi1, phi2 = math.radians(lat1), math.radians(lat2)
    dphi = math.radians(lat2 - lat1)
    dlambda = math.radians(lon2 - lon1)
    
    a = math.sin(dphi/2)**2 + \
        math.cos(phi1)*math.cos(phi2)*math.sin(dlambda/2)**2
    
    return 2*R*math.atan2(math.sqrt(a), math.sqrt(1 - a))

# 按行程ID分组,获取每个点的下一个点经纬度
df['next_lat'] = df.groupby('tripId')['latitude'].shift(-1)
df['next_lon'] = df.groupby('tripId')['longitude'].shift(-1)

# 逐行计算两点距离
df['Distance'] = df.apply(lambda row: haversine(
    (row['latitude'], row['longitude']), 
    (row['next_lat'], row['next_lon'])
), axis=1)

# 清理临时字段
df.drop(columns=['next_lat', 'next_lon'], inplace=True)

效果说明

执行后每个tripId分组的最后一行Distance值为空,有效距离值的数量正好是分组行数减1,符合你的预期。

内容的提问来源于stack exchange,提问作者fifi.h

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 01:18:05