如何在Pandas中基于经纬度列批量计算生成距离新列?
批量计算经纬度距离生成新列
前置条件
- 已导入
haversine库:import haversine as hs - 需要补充导入
numpy库:import numpy as np(函数中用到np.nan) - 数据存储在Pandas DataFrame中,包含两列经纬度元组:
| Customer_lat_lon | branch_lat_lon |
|---|---|
| (28.682,77.175) | (28.599,77.334) |
| (19.126,72.865) | (19.104,72.863) |
修正计算函数
原函数存在语法错误(缺少参数定义),正确版本如下:
def calc_distance(x, y): try: return hs.haversine(x, y) except: return np.nan
批量生成distance列
使用Pandas的apply方法逐行调用函数,生成新列:
# 假设数据框名为df df['distance'] = df.apply(lambda row: calc_distance(row['Customer_lat_lon'], row['branch_lat_lon']), axis=1)
结果示例
执行后的数据框会新增distance列,示例如下:
| Customer_lat_lon | branch_lat_lon | distance |
|---|---|---|
| (28.682,77.175) | (28.599,77.334) | 18.0612 |
| (19.126,72.865) | (19.104,72.863) | 2.4578 |
说明
- 该方法可高效处理1000条记录
- 异常捕获机制确保遇到无效经纬度时返回
NaN,避免程序中断
内容的提问来源于stack exchange,提问作者Sasi
相关产品推荐
相关产品推荐

