You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于注意力的图像字幕模型.fit()报错:train_function返回空日志

问题:基于注意力机制的图像字幕模型训练时train_step未触发,导致日志为空错误

问题描述

实现基于注意力机制的图像字幕编解码器模型后,调用model.fit()时触发以下错误:

ValueError: Unexpected result of train_function (Empty logs). Please use Model.compile(..., run_eagerly=True), or tf.config.run_functions_eagerly(True) for more information of where went wrong, or file a issue/bug to tf.keras.

尝试设置run_eagerly=True和tf.config.run_functions_eagerly(True)后问题仍未解决,调试发现:

  • call方法返回的batch_loss是正常的Tensor值(如<tf.Tensor: shape=(), dtype=float32, numpy=2.7348917>)
  • train_step方法的断点从未触发,fit()直接调用了call方法而非train_step

模型代码

import tensorflow as tf
from ..layers.encoder import Encoder
from ..layers.decoder import Decoder


class CaptionerModel(tf.keras.Model):

    def __init__(self, tokenizer_wrapper, embedding_dim=256, units=512):
        """
        Constructor method for the CaptionerModel class.
        :param units: The number of units to use in the GRU layer
        """

        tf.config.run_functions_eagerly(True)

        super(CaptionerModel, self).__init__()
        self.tokenizer = tokenizer_wrapper.tokenizer
        self.vocab_size = len(self.tokenizer.word_index) + 1
        self.max_length = tokenizer_wrapper.max_length
        self.embedding_dim = embedding_dim
        self.units = units

        # encoder layers
        self.encoder = Encoder(embedding_dim)

        # decoder layer
        self.decoder = Decoder(self.embedding_dim, self.units, self.vocab_size)

    def call(self, data):
        """
        This method performs the forward pass of the model.
        :param images: Input images
        :param importance_features: Importance features for the images
        :param captions: Target captions
        :return: Predicted captions
        """

        # get the images, importance features and captions
        images, importance_features, captions = data

        # initialise the loss
        loss = 0

        # initialise the hidden state
        hidden = self.decoder.reset_state(batch_size=images.shape[0])

        # create the decoder input (which is the start token)
        dec_input = tf.expand_dims([self.tokenizer.word_index["<start>"]] * images.shape[0], 1)

        # get the image features
        features = self.encoder(images, importance_features)

        # iterate over the captions
        for i in range(1, captions.shape[1]):

            # get the predictions
            predictions, hidden, _ = self.decoder(dec_input, features, hidden)

            # calculate the loss
            loss += self.loss_function(captions[:, i], predictions)

            # use teacher forcing
            dec_input = tf.expand_dims(captions[:, i], 1)

        # calculate the batch loss
        batch_loss = (loss / int(captions.shape[1]))

        return batch_loss

    def train_step(self, data):
        """
        This method performs a single training step.
        :param data: The data to train on in the form of a tuple of (images, importance_features, captions)
        """
        images, importance_features, captions = data

        with tf.GradientTape() as tape:
            l = self((images, importance_features, captions))

        # get the trainable variables
        trainable_variables = self.encoder.trainable_variables + self.decoder.trainable_variables

        # calculate the gradients
        gradients = tape.gradient(l, trainable_variables)

        # apply the gradients
        self.optimizer.apply_gradients(zip(gradients, trainable_variables))

        # update the metrics
        self.loss(l)

        return {"loss": self.loss.result()}
    
    def loss_function(self, real, pred):
        """
        This method calculates the loss.
        """
        mask = tf.math.logical_not(tf.math.equal(real, 0))
        loss_ = tf.keras.losses.sparse_categorical_crossentropy(real, pred, from_logits=True)

        mask = tf.cast(mask, dtype=loss_.dtype)
        loss_ *= mask

        return tf.reduce_mean(loss_)

错误信息

ValueError: Unexpected result of `train_function` (Empty logs). Please use `Model.compile(..., run_eagerly=True)`, or `tf.config.run_functions_eagerly(True)` for more information of where went wrong, or file a issue/bug to `tf.keras`.

解决方案

核心问题是**call方法的返回值不符合Keras Model的设计预期**:

  • 自定义train_step时,call方法应返回模型的预测结果(字幕的概率分布),而非直接返回损失值。Keras会在train_step中调用call获取预测,再完成损失计算和梯度更新。
  • 当前call直接返回损失,打乱了Keras的训练流程,导致train_step被跳过。

修改步骤如下:

1. 重构call方法,返回预测序列

让call根据训练/推理模式返回对应的预测结果,将损失计算逻辑移到train_step中:

def call(self, data, training=False):
    # 区分训练和推理模式的输入
    if training:
        images, importance_features, captions = data
    else:
        images, importance_features = data
        captions = None
    
    features = self.encoder(images, importance_features)
    hidden = self.decoder.reset_state(batch_size=images.shape[0])
    dec_input = tf.expand_dims([self.tokenizer.word_index["<start>"]] * images.shape[0], 1)
    predictions = []
    
    if training:
        # 训练阶段使用teacher forcing
        for i in range(1, captions.shape[1]):
            pred, hidden, _ = self.decoder(dec_input, features, hidden)
            predictions.append(pred)
            dec_input = tf.expand_dims(captions[:, i], 1)
    else:
        # 推理阶段生成完整字幕
        for i in range(self.max_length):
            pred, hidden, _ = self.decoder(dec_input, features, hidden)
            predicted_id = tf.argmax(pred, axis=1)
            predictions.append(pred)
            dec_input = tf.expand_dims(predicted_id, 1)
            if predicted_id == self.tokenizer.word_index["<end>"]:
                break
    
    # 将预测序列转为(batch_size, max_length, vocab_size)格式的张量
    return tf.stack(predictions, axis=1)

2. 调整train_step的损失计算逻辑

调用重构后的call获取预测,再用自定义损失函数计算损失:

def train_step(self, data):
    images, importance_features, captions = data

    with tf.GradientTape() as tape:
        predictions = self((images, importance_features, captions), training=True)
        # 计算损失,注意跳过caption的<start> token
        loss = 0
        for i in range(predictions.shape[1]):
            loss += self.loss_function(captions[:, i+1], predictions[:, i])
        batch_loss = loss / int(predictions.shape[1])

    # 梯度更新流程
    trainable_variables = self.encoder.trainable_variables + self.decoder.trainable_variables
    gradients = tape.gradient(batch_loss, trainable_variables)
    self.optimizer.apply_gradients(zip(gradients, trainable_variables))

    # 更新损失指标
    self.loss(batch_loss)
    return {"loss": self.loss.result()}

3. 编译模型时的兼容设置

编译时可显式指定损失函数(实际损失在train_step中计算,这里仅做兼容):

model.compile(optimizer=tf.keras.optimizers.Adam(), loss=lambda y_true, y_pred: y_pred)

完成以上修改后,train_step会被正常调用,训练流程恢复正常,日志为空的错误也会解决。


内容的提问来源于stack exchange,提问作者Omar El Atyqy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 20:55:55