模型拟合期间如何调试自定义损失函数?
问题
希望查看模型拟合过程中自定义损失函数内的运行情况,但尝试以下代码未成功:
def custom_loss(label : tf.Tensor, pred : tf.Tensor) -> tf.Tensor: mask = label != 0 loss_object = tf.keras.losses.BinaryCrossentropy(from_logits=True, reduction='none') loss = loss_object(label, pred) mask = tf.cast(mask, dtype=loss.dtype) tf.print("\n---------------------------") tf.print("custom_loss - str(loss):", str(loss)) tf.print("custom_loss - str(mask):", str(mask)) try: tf.print("tf.keras.backend.eval(loss):", tf.keras.backend.eval(loss)) except: tf.print("tf.keras.backend.eval(loss) does not work - exception!") loss = tf.reshape(loss, shape=(batch_size, loss.shape[1], 1)) loss *= mask loss = tf.reduce_sum(loss)/tf.reduce_sum(mask) tf.print("\n============================") return loss
调用fit()启动训练后,仅得到以下输出:
2/277 [..............................] - ETA: 44s - loss: 0.6931 - masked_accuracy: 0.0000e+00 --------------------------- custom_loss - str(loss): Tensor("custom_loss/binary_crossentropy/weighted_loss/Mul:0", shape=(None, 20), dtype=float32) custom_loss - str(mask): Tensor("custom_loss/Cast:0", shape=(None, 20, 1), dtype=float32) tf.keras.backend.eval(loss) does not work - exception!
请问如何显示label、pred、mask和loss的实际值?
解决方案
要在自定义损失函数中打印张量的实际数值,核心问题是TensorFlow图模式下,str()只能输出张量的符号化信息,tf.keras.backend.eval()无法直接在图中执行。可以用以下几种方法解决:
直接用
tf.print()打印张量本身
不要用str()包裹张量,直接将张量传入tf.print(),它会在图执行阶段输出张量的实际数值。修改后的代码片段如下:tf.print("\n---------------------------") tf.print("custom_loss - label:", label) tf.print("custom_loss - pred:", pred) tf.print("custom_loss - loss:", loss) tf.print("custom_loss - mask:", mask)这样就能输出每个批次中这些张量的具体数值。
控制打印频率(避免刷屏)
如果不想每个批次都打印,可以添加条件判断,只打印前N个批次的信息,同时注意不要用固定batch_size,改用动态维度获取避免错误:# 定义变量记录批次次数 batch_count = tf.Variable(0, dtype=tf.int32) def custom_loss(label : tf.Tensor, pred : tf.Tensor) -> tf.Tensor: batch_count.assign_add(1) mask = label != 0 loss_object = tf.keras.losses.BinaryCrossentropy(from_logits=True, reduction='none') loss = loss_object(label, pred) mask = tf.cast(mask, dtype=loss.dtype) # 仅打印前3个批次的信息 tf.cond(batch_count <= 3, lambda: tf.print("\n---------------------------", "\ncustom_loss - label:", label, "\ncustom_loss - pred:", pred, "\ncustom_loss - loss:", loss, "\ncustom_loss - mask:", mask), lambda: tf.no_op()) # 用动态维度替代固定batch_size loss = tf.reshape(loss, shape=(tf.shape(loss)[0], loss.shape[1], 1)) loss *= mask loss = tf.reduce_sum(loss)/tf.reduce_sum(mask) tf.cond(batch_count <= 3, lambda: tf.print("\n============================", "final loss:", loss), lambda: tf.no_op()) return loss切换到即刻执行模式(仅调试用)
临时关闭图模式,切换到即刻执行,这样可以用普通print()直接打印张量的numpy值,但会降低训练速度,仅适合调试阶段:# 开启即刻执行 tf.config.run_functions_eagerly(True) # 启动训练 model.fit(...) # 调试完成后切回图模式 tf.config.run_functions_eagerly(False)此时在损失函数中可以直接写:
print("custom_loss - label:", label.numpy()) print("custom_loss - pred:", pred.numpy())
内容的提问来源于stack exchange,提问作者toom
相关产品推荐
相关产品推荐

