You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义Haversine Loss函数报错:无梯度提供,求解决方案

问题描述

我想实现一个自定义损失函数——Haversine Loss,用来计算球面上两点间的距离。类的坐标存储在CSV文件中,需要将真实标签和预测结果的索引转换为类名('cluster'),再从CSV中提取对应坐标(CSV仅包含cluster整数、Center_Longitude和Center_Latitude浮点型数据)。坐标提取与Haversine距离计算功能正常(见下方输出),但最终报错:ValueError: No gradients provided for any variable。

我的模型输出层为softmax,因此使用argmax获取概率最高的预测结果,我怀疑这是问题根源,但不确定如何解决。请问:

  1. 为何会出现无梯度的情况?
  2. 如何修改代码使其正常运行?

自定义损失函数代码

def haversine_loss(y_true, y_pred):

    pred_index = tf.argmax(y_pred, axis=-1)
    true_index = tf.cast(y_true, tf.int32)


    tf.print("Predicted indices:", pred_index)
    tf.print("True indices:", true_index)

   
    pred_class_names = tf.gather(class_names, pred_index)
    true_class_names = tf.gather(class_names, true_index)

    
    tf.print("Predicted class names:", pred_class_names)
    tf.print("True class names:", true_class_names)


    pred_lat_lon = tf.stack([
        tf.gather(clusters_df['Center_Latitude'].values, pred_index),
        tf.gather(clusters_df['Center_Longitude'].values, pred_index)
    ], axis=-1)

    true_lat_lon = tf.stack([
        tf.gather(clusters_df['Center_Latitude'].values, true_index),
        tf.gather(clusters_df['Center_Longitude'].values, true_index)
    ], axis=-1)

  
    tf.print("Predicted lat/lon:", pred_lat_lon)
    tf.print("True lat/lon:", true_lat_lon)

 
    lat1, lon1 = pred_lat_lon[:, 0], pred_lat_lon[:, 1]
    lat2, lon2 = true_lat_lon[:, 0], true_lat_lon[:, 1]

  
    distance = haversine_distance(lat1, lon1, lat2, lon2)

    tf.print("Calculated distances:", distance)

    return tf.reduce_mean(distance)

Haversine距离计算函数

def haversine_distance(lat1, lon1, lat2, lon2):

    lat1 = tf.cast(lat1, tf.float32) * (tf.constant(m.pi, dtype=tf.float32) / 180.0)
    lon1 = tf.cast(lon1, tf.float32) * (tf.constant(m.pi, dtype=tf.float32) / 180.0)
    lat2 = tf.cast(lat2, tf.float32) * (tf.constant(m.pi, dtype=tf.float32) / 180.0)
    lon2 = tf.cast(lon2, tf.float32) * (tf.constant(m.pi, dtype=tf.float32) / 180.0)

    dlon = lon2 - lon1
    dlat = lat2 - lat1

    a = tf.sin(dlat / 2)**2 + tf.cos(lat1) * tf.cos(lat2) * tf.sin(dlon / 2)**2
    c = 2 * tf.asin(tf.sqrt(a))

    R = 6371.0
    return R * c

报错信息

Epoch 1/25
Predicted indices: [1124]
True indices: [604]
Predicted class names: ["8983"]
True class names: ["8463"]
Predicted lat/lon: [[51.227374672500062 7.2118920326250366]]
True lat/lon: [[49.075162783850026 9.48693443665003]]
Calculated distances: [289.012787]
Traceback (most recent call last):
  File "mypath.py", line 171, in <module>
    model.fit(
  File "mypath.py\pythonProject\.venv\Lib\site-packages\keras\src\utils\traceback_utils.py", line 122, in error_handler
    raise e.with_traceback(filtered_tb) from None
  File "mypath.py\pythonProject\.venv\Lib\site-packages\keras\src\optimizers\base_optimizer.py", line 662, in _filter_empty_gradients
    raise ValueError("No gradients provided for any variable.")
ValueError: No gradients provided for any variable.

Process finished with exit code 1

问题原因与解决方案

原因分析

核心问题确实是tf.argmax(y_pred, axis=-1):argmax是不可导操作,它直接从softmax输出的概率分布中选取最大概率的索引,这个离散选择过程会彻底切断梯度传递路径——模型参数无法通过这个操作获得更新信号,最终导致优化器找不到可更新的变量梯度,抛出错误。此外,基于索引的tf.gather操作也会进一步阻断梯度,但根源还是argmax的使用。

解决方案

不能将softmax输出的连续概率分布转为离散索引,而是要保留概率分布,通过加权求和计算期望坐标,再基于期望坐标计算损失。具体步骤如下:

  1. 预处理CSV数据,将所有cluster的坐标整理为形状为[类别数, 2]的Tensor矩阵,每行对应一个cluster的纬度和经度。
  2. 用模型输出的softmax概率(y_pred)与坐标矩阵做矩阵乘法,得到预测坐标的加权平均值(期望坐标)。
  3. 真实标签对应的坐标直接用tf.gather从坐标矩阵中提取(真实标签为固定值,无需计算梯度)。
  4. 用Haversine距离计算期望预测坐标与真实坐标的距离,作为最终损失。

修改后的代码示例

1. 预处理坐标矩阵

# 将CSV中的坐标转换为Tensor矩阵(确保cluster索引与模型输出类别索引对应)
coords_matrix = tf.convert_to_tensor(
    clusters_df[['Center_Latitude', 'Center_Longitude']].values,
    dtype=tf.float32
)

# 如果cluster编号与矩阵索引不匹配,需先做映射:
# cluster_to_idx = {cluster: idx for idx, cluster in enumerate(clusters_df['cluster'].values)}
# 真实标签需先转换为对应索引再传入损失函数

2. 修改损失函数

def haversine_loss(y_true, y_pred):
    # y_true:真实cluster索引(int32);y_pred:softmax输出的概率分布(shape [batch_size, num_classes])
    
    # 计算预测的期望坐标:概率加权求和
    pred_lat_lon = tf.matmul(y_pred, coords_matrix)  # shape [batch_size, 2]
    
    # 获取真实坐标
    true_lat_lon = tf.gather(coords_matrix, tf.cast(y_true, tf.int32))  # shape [batch_size, 2]
    
    # 计算Haversine距离
    lat1, lon1 = pred_lat_lon[:, 0], pred_lat_lon[:, 1]
    lat2, lon2 = true_lat_lon[:, 0], true_lat_lon[:, 1]
    distance = haversine_distance(lat1, lon1, lat2, lon2)
    
    return tf.reduce_mean(distance)

关键说明

  • 保留softmax概率分布并计算期望坐标,整个过程连续可导,梯度能正常传递回模型参数。
  • 真实标签的坐标提取用tf.gather不影响梯度计算,因为真实标签是固定值,无需对其求导。
  • 这种方式让模型学习输出一个概率分布,使该分布对应的期望坐标尽可能接近真实坐标,更贴合距离损失的优化目标。

内容的提问来源于stack exchange,提问作者d3vilstrap

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 12:25:55