You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

定义的Haversine函数未正确遍历数据,生成站点距离矩阵全触发异常

站点距离矩阵全为默认异常值问题排查与解决

问题背景

尝试生成站点间距离的二维矩阵,已实现distance函数计算两点间哈维正弦距离,distance_harvesine函数遍历DataFrame生成矩阵,但调用后所有值均为异常分支的默认值100000000000,无法得到正确距离。

现有代码

距离计算函数

import math
import pandas as pd

def distance(origin, destination):
    lat1, lon1 = origin
    lat2, lon2 = destination
    radius = 6371 # km

    dlat = math.radians(lat2-lat1)
    dlon = math.radians(lon2-lon1)
    a = math.sin(dlat/2) * math.sin(dlat/2) + math.cos(math.radians(lat1)) \
        * math.cos(math.radians(lat2)) * math.sin(dlon/2) * math.sin(dlon/2)
    c = 2 * math.atan2(math.sqrt(a), math.sqrt(1-a))
    d = radius * c

    return d

矩阵生成函数

def distance_harvesine(df):
    # 使用字典存储,无需预定义大小
    dict_distance = {}
    
    df_copy = df.copy()
    
    for i, row in df_copy.iterrows():
        list_distance = []
        
        lon1 = df_copy['longitude'][i]
        lat1 = df_copy['latitude'][i]
        
        origin = (lat1,lon1)

        for j, row in df_copy.iterrows():
            lon2 = df_copy['longitude'][j]
            lat2 = df_copy['latitude'][j]
            
            destination = (lat2,lon2)
            
            try:
                dist = distance(origin, destination)
            except:
                dist = 100000000000
                                
            list_distance.append(dist)
            
        dict_distance[i] = list_distance
            

    distance_matrix = pd.DataFrame.from_dict(dict_distance)

    return distance_matrix

调用代码

df_H = distance_harvesine(df)

问题原因分析

  1. 异常捕获无日志:原代码的try-except块捕获所有异常但未输出错误信息,无法定位具体触发异常的原因。
  2. 数据类型/空值问题:大概率是DataFrame中的latitude或longitude列存在非数值类型(如字符串)、NaN值,导致math模块的函数(如math.radians)抛出异常。
  3. 取值效率问题:通过df_copy['longitude'][i]的方式取值效率较低,但不是导致异常的直接原因。

解决方案

1. 排查并修复数据问题

先检查经纬度列的数据类型和空值情况:

# 查看数据类型
print(df[['latitude', 'longitude']].dtypes)
# 统计空值数量
print(df[['latitude', 'longitude']].isna().sum())

如果是字符串类型,转换为数值并处理空值:

# 转换为数值类型,无法转换的设为NaN
df['latitude'] = pd.to_numeric(df['latitude'], errors='coerce')
df['longitude'] = pd.to_numeric(df['longitude'], errors='coerce')
# 删除包含空值的行
df = df.dropna(subset=['latitude', 'longitude'])

2. 修改异常捕获,添加错误日志

修改try-except块,打印具体异常信息,方便定位问题:

try:
    dist = distance(origin, destination)
except Exception as e:
    print(f"计算原点{origin}到目标点{destination}时出错: {str(e)}")
    dist = 100000000000

3. 优化遍历逻辑(可选)

直接使用迭代行中的值,提升代码效率:

def distance_harvesine(df):
    dict_distance = {}
    # 重置索引避免迭代时出现索引混乱
    df_copy = df.copy().reset_index(drop=True)
    
    for i, row_i in df_copy.iterrows():
        list_distance = []
        origin = (row_i['latitude'], row_i['longitude'])
        
        for j, row_j in df_copy.iterrows():
            destination = (row_j['latitude'], row_j['longitude'])
            try:
                dist = distance(origin, destination)
            except Exception as e:
                print(f"索引{i}到{j}的距离计算出错: {str(e)}")
                dist = 100000000000
            list_distance.append(dist)
            
        dict_distance[i] = list_distance
    
    return pd.DataFrame.from_dict(dict_distance)

内容的提问来源于stack exchange,提问作者Marco Rous

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 14:45:57