基于注意力的图像字幕模型.fit()报错:train_function返回空日志
问题:基于注意力机制的图像字幕模型训练时train_step未触发,导致日志为空错误
问题描述
实现基于注意力机制的图像字幕编解码器模型后,调用model.fit()时触发以下错误:
ValueError: Unexpected result of
train_function(Empty logs). Please useModel.compile(..., run_eagerly=True), ortf.config.run_functions_eagerly(True)for more information of where went wrong, or file a issue/bug totf.keras.
尝试设置run_eagerly=True和tf.config.run_functions_eagerly(True)后问题仍未解决,调试发现:
call方法返回的batch_loss是正常的Tensor值(如<tf.Tensor: shape=(), dtype=float32, numpy=2.7348917>)train_step方法的断点从未触发,fit()直接调用了call方法而非train_step
模型代码
import tensorflow as tf from ..layers.encoder import Encoder from ..layers.decoder import Decoder class CaptionerModel(tf.keras.Model): def __init__(self, tokenizer_wrapper, embedding_dim=256, units=512): """ Constructor method for the CaptionerModel class. :param units: The number of units to use in the GRU layer """ tf.config.run_functions_eagerly(True) super(CaptionerModel, self).__init__() self.tokenizer = tokenizer_wrapper.tokenizer self.vocab_size = len(self.tokenizer.word_index) + 1 self.max_length = tokenizer_wrapper.max_length self.embedding_dim = embedding_dim self.units = units # encoder layers self.encoder = Encoder(embedding_dim) # decoder layer self.decoder = Decoder(self.embedding_dim, self.units, self.vocab_size) def call(self, data): """ This method performs the forward pass of the model. :param images: Input images :param importance_features: Importance features for the images :param captions: Target captions :return: Predicted captions """ # get the images, importance features and captions images, importance_features, captions = data # initialise the loss loss = 0 # initialise the hidden state hidden = self.decoder.reset_state(batch_size=images.shape[0]) # create the decoder input (which is the start token) dec_input = tf.expand_dims([self.tokenizer.word_index["<start>"]] * images.shape[0], 1) # get the image features features = self.encoder(images, importance_features) # iterate over the captions for i in range(1, captions.shape[1]): # get the predictions predictions, hidden, _ = self.decoder(dec_input, features, hidden) # calculate the loss loss += self.loss_function(captions[:, i], predictions) # use teacher forcing dec_input = tf.expand_dims(captions[:, i], 1) # calculate the batch loss batch_loss = (loss / int(captions.shape[1])) return batch_loss def train_step(self, data): """ This method performs a single training step. :param data: The data to train on in the form of a tuple of (images, importance_features, captions) """ images, importance_features, captions = data with tf.GradientTape() as tape: l = self((images, importance_features, captions)) # get the trainable variables trainable_variables = self.encoder.trainable_variables + self.decoder.trainable_variables # calculate the gradients gradients = tape.gradient(l, trainable_variables) # apply the gradients self.optimizer.apply_gradients(zip(gradients, trainable_variables)) # update the metrics self.loss(l) return {"loss": self.loss.result()} def loss_function(self, real, pred): """ This method calculates the loss. """ mask = tf.math.logical_not(tf.math.equal(real, 0)) loss_ = tf.keras.losses.sparse_categorical_crossentropy(real, pred, from_logits=True) mask = tf.cast(mask, dtype=loss_.dtype) loss_ *= mask return tf.reduce_mean(loss_)
错误信息
ValueError: Unexpected result of `train_function` (Empty logs). Please use `Model.compile(..., run_eagerly=True)`, or `tf.config.run_functions_eagerly(True)` for more information of where went wrong, or file a issue/bug to `tf.keras`.
解决方案
核心问题是**call方法的返回值不符合Keras Model的设计预期**:
- 自定义
train_step时,call方法应返回模型的预测结果(字幕的概率分布),而非直接返回损失值。Keras会在train_step中调用call获取预测,再完成损失计算和梯度更新。 - 当前
call直接返回损失,打乱了Keras的训练流程,导致train_step被跳过。
修改步骤如下:
1. 重构call方法,返回预测序列
让call根据训练/推理模式返回对应的预测结果,将损失计算逻辑移到train_step中:
def call(self, data, training=False): # 区分训练和推理模式的输入 if training: images, importance_features, captions = data else: images, importance_features = data captions = None features = self.encoder(images, importance_features) hidden = self.decoder.reset_state(batch_size=images.shape[0]) dec_input = tf.expand_dims([self.tokenizer.word_index["<start>"]] * images.shape[0], 1) predictions = [] if training: # 训练阶段使用teacher forcing for i in range(1, captions.shape[1]): pred, hidden, _ = self.decoder(dec_input, features, hidden) predictions.append(pred) dec_input = tf.expand_dims(captions[:, i], 1) else: # 推理阶段生成完整字幕 for i in range(self.max_length): pred, hidden, _ = self.decoder(dec_input, features, hidden) predicted_id = tf.argmax(pred, axis=1) predictions.append(pred) dec_input = tf.expand_dims(predicted_id, 1) if predicted_id == self.tokenizer.word_index["<end>"]: break # 将预测序列转为(batch_size, max_length, vocab_size)格式的张量 return tf.stack(predictions, axis=1)
2. 调整train_step的损失计算逻辑
调用重构后的call获取预测,再用自定义损失函数计算损失:
def train_step(self, data): images, importance_features, captions = data with tf.GradientTape() as tape: predictions = self((images, importance_features, captions), training=True) # 计算损失,注意跳过caption的<start> token loss = 0 for i in range(predictions.shape[1]): loss += self.loss_function(captions[:, i+1], predictions[:, i]) batch_loss = loss / int(predictions.shape[1]) # 梯度更新流程 trainable_variables = self.encoder.trainable_variables + self.decoder.trainable_variables gradients = tape.gradient(batch_loss, trainable_variables) self.optimizer.apply_gradients(zip(gradients, trainable_variables)) # 更新损失指标 self.loss(batch_loss) return {"loss": self.loss.result()}
3. 编译模型时的兼容设置
编译时可显式指定损失函数(实际损失在train_step中计算,这里仅做兼容):
model.compile(optimizer=tf.keras.optimizers.Adam(), loss=lambda y_true, y_pred: y_pred)
完成以上修改后,train_step会被正常调用,训练流程恢复正常,日志为空的错误也会解决。
内容的提问来源于stack exchange,提问作者Omar El Atyqy
相关产品推荐
相关产品推荐

