You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为LSTM实现基于惩罚矩阵的Keras自定义损失函数

问题背景

我正在处理多分类任务,采用LSTM模型实现,此前一直使用categorical_crossentropy作为损失函数训练模型。训练完成后校验模型效果时,使用以下自定义指标计算得分,其中A为12×12的二维惩罚矩阵,与12个分类维度对应:

def score(y_true, y_pred):
    S = 0.0
    y_true = y_true.astype(int)
    y_pred = y_pred.astype(int)
    for i in range(0, y_true.shape[0]):
        S -= A[y_true[i], y_pred[i]]
    return S/y_true.shape[0]

该自定义指标支持输入Pandas Series类型的y_true和y_pred,输出为负数,数值越接近0代表模型效果越好。
我现在想要把当前使用的categorical_crossentropy损失函数,替换为和上述自定义指标逻辑一致、可以引入A惩罚矩阵计算的自定义损失。

遇到的问题

损失函数的输入是Tensor类型而非Pandas Series对象,且因为使用LSTM,输入的Tensor为3维结构,参数如下:

y_true: Tensor("IteratorGetNext:1", shape=(1, 25131, 12), dtype=uint8)
type(y_true): <class 'tensorflow.python.framework.ops.Tensor'>
y_pred: Tensor("sequential_26/time_distributed_26/Reshape_1:0", shape=(1, 25131, 12), dtype=float32)
type(y_pred): <class 'tensorflow.python.framework.ops.Tensor'>
模型架构与数据维度

我的模型结构如下:

callbacks = [EarlyStopping(monitor='val_loss', patience=25)]

model = Sequential()
model.add(Masking(mask_value = 0.))
model.add(Bidirectional(LSTM(64, return_sequences=True, activation = "tanh")))
model.add(Dropout(0.3))
model.add(TimeDistributed(Dense(12, activation='softmax')))
adam = adam_v2.Adam(learning_rate=0.002)

model.compile(optimizer=adam, loss=score, metrics=['accuracy'])

history = model.fit(X_train, y_train, epochs=150, batch_size=1, shuffle=False,
                    validation_data=(X_test, y_test), verbose=2, callbacks=[callbacks])

模型输入输出的维度如下,共12个分类:

print(f'{X_train.shape} {X_test.shape} {y_train.shape} {y_test.shape}')
# 输出:(73, 25131, 29) (25, 23879, 29) (73, 25131, 12) (25, 23879, 12)

A为12×12的惩罚矩阵,维度与多分类任务的分类数对应:
惩罚矩阵示意图

该模型为岩性识别竞赛搭建。


解决方案

要实现适配TensorFlow张量、支持3维时序输入的自定义损失,按以下步骤调整即可:

  1. 首先把惩罚矩阵A转换为TensorFlow常量,保证可以在计算图中正常调用:
import tensorflow as tf
# 假设A是你已经定义好的12*12 numpy数组
A_tensor = tf.constant(A, dtype=tf.float32)
  1. 重写自定义损失函数,替换numpy算子为TensorFlow原生算子,同时适配3维输入和Masking逻辑:
def custom_loss(y_true, y_pred):
    # 把独热编码的y_true转成类别索引,shape从(batch, 时序长度, 12)转为(batch, 时序长度)
    y_true_idx = tf.argmax(y_true, axis=-1, output_type=tf.int32)
    y_pred_idx = tf.argmax(y_pred, axis=-1, output_type=tf.int32)
    # 拼接索引得到每个样本点对应A矩阵的坐标
    indices = tf.stack([y_true_idx, y_pred_idx], axis=-1)
    # 取出对应惩罚值,取负数和原指标逻辑对齐
    penalty = -tf.gather_nd(A_tensor, indices)
    # 匹配Masking层逻辑,过滤全0的填充位置
    mask = tf.cast(tf.reduce_any(y_true != 0, axis=-1), dtype=tf.float32)
    # 仅统计非填充位置的平均损失
    loss = tf.reduce_sum(penalty * mask) / tf.reduce_sum(mask)
    return loss
  1. 替换模型编译时的损失函数即可:
model.compile(optimizer=adam, loss=custom_loss, metrics=['accuracy'])

优化建议

上面的损失直接使用argmax取预测类别,但argmax是不可导算子,可能导致训练梯度中断,收敛不稳定。更推荐使用带惩罚权重的交叉熵实现,既可以保留梯度,又能贴合你的惩罚逻辑:

def weighted_categorical_crossentropy(y_true, y_pred):
    y_true_idx = tf.argmax(y_true, axis=-1, output_type=tf.int32)
    # 取出真实类别对应的惩罚权重行
    weights = tf.gather(A_tensor, y_true_idx)
    # 计算带惩罚权重的交叉熵
    cce = tf.keras.losses.categorical_crossentropy(y_true, y_pred)
    weighted_cce = cce * tf.reduce_sum(y_true * weights, axis=-1)
    # 同样加入Mask逻辑
    mask = tf.cast(tf.reduce_any(y_true != 0, axis=-1), dtype=tf.float32)
    loss = tf.reduce_sum(weighted_cce * mask) / tf.reduce_sum(mask)
    return loss

该实现训练收敛性更好,最终的指标表现也会更贴合你要优化的惩罚得分。


内容的提问来源于stack exchange,提问作者Matheus Schaly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 04:24:01