You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于InceptionResNetV2的三元组模型训练梯度缺失错误求助

问题分析与解决方案

核心错误原因

  1. 损失函数参数顺序颠倒:Keras自定义损失函数的标准签名是def loss(y_true, y_pred, ...),你将y_pred放在了第一个参数位置,导致损失计算时输入匹配错误,梯度无法关联到可训练变量。
  2. 生成器输出与模型输出不匹配:你的FRmodel输出3个特征向量,但生成器返回的是单个dummy张量,Keras无法正确关联损失与模型输出。
  3. 距离计算轴参数错误:使用axis=0会跨样本维度求和,而非对单个样本的特征维度计算距离,破坏了梯度传播路径。

具体修复步骤

1. 修正损失函数

调整参数顺序,并修正距离计算的轴参数:

def triplet_loss(y_true, y_pred, alpha=0.1):
    anchor, positive, negative = y_pred[0], y_pred[1], y_pred[2]

    # 对单个样本的特征维度(axis=1)计算欧氏距离平方
    pos_dist = tf.reduce_sum(tf.square(tf.subtract(anchor, positive)), axis=1)
    neg_dist = tf.reduce_sum(tf.square(tf.subtract(anchor, negative)), axis=1)
        
    basic_loss = tf.add(tf.subtract(pos_dist, neg_dist), alpha)
    loss = tf.reduce_mean(tf.maximum(basic_loss, 0.0))
      
    return loss

2. 调整生成器输出格式

生成与模型输出数量匹配的dummy张量,同时修正索引越界问题:

def data_generator(batch_size, name='train'):
    # 生成3个dummy张量,对应模型的3个输出
    dummy_target = [np.zeros((batch_size, 128)) for _ in range(3)]
    anchors = np.zeros((batch_size, 218, 178, 3))
    positives = np.zeros((batch_size, 218, 178, 3))
    negatives = np.zeros((batch_size, 218, 178, 3))
    while True:
        included_files_in_batch = []
        for i in range(batch_size):
            # 修正索引越界:randint上限设为len-1
            file_num = random.randint(0, len(train_file_list)-1)
            while file_num in included_files_in_batch:
                file_num = random.randint(0, len(train_file_list)-1)
            if name == 'train':
                file = train_file_list[file_num]
            else:
                file = val_file_list[file_num]
            
            anchors[i] = load_file(file, 'anchor')
            positives[i] = load_file(file, 'positive')
            negatives[i] = load_file(file, 'negative')
            
            included_files_in_batch.append(file_num)
            
        data = {
                'anchors': anchors,
                'positives': positives,
                'negatives': negatives
            }
            
        yield (data, dummy_target)

3. 模型编译配置

直接使用修正后的损失函数编译模型:

FRmodel.compile(optimizer='adam', loss=triplet_loss)

4. 额外验证点

  • 确认initial_base_model.trainable = False是预期设置(仅训练顶部Dense层),后续如需微调基础模型可再开启,当前错误与该设置无关。
  • tf.math.l2_normalize是可微分操作,无需移除,不影响梯度传播。

修复效果验证

完成上述修改后重新训练,梯度将正确传播到dense_layer的参数,不会再出现"No gradients provided"的错误。

内容的提问来源于stack exchange,提问作者Extra_Caterpillar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 22:55:02