如何为LSTM实现基于惩罚矩阵的Keras自定义损失函数
问题背景
我正在处理多分类任务,采用LSTM模型实现,此前一直使用categorical_crossentropy作为损失函数训练模型。训练完成后校验模型效果时,使用以下自定义指标计算得分,其中A为12×12的二维惩罚矩阵,与12个分类维度对应:
def score(y_true, y_pred): S = 0.0 y_true = y_true.astype(int) y_pred = y_pred.astype(int) for i in range(0, y_true.shape[0]): S -= A[y_true[i], y_pred[i]] return S/y_true.shape[0]
该自定义指标支持输入Pandas Series类型的y_true和y_pred,输出为负数,数值越接近0代表模型效果越好。
我现在想要把当前使用的categorical_crossentropy损失函数,替换为和上述自定义指标逻辑一致、可以引入A惩罚矩阵计算的自定义损失。
遇到的问题
损失函数的输入是Tensor类型而非Pandas Series对象,且因为使用LSTM,输入的Tensor为3维结构,参数如下:
y_true: Tensor("IteratorGetNext:1", shape=(1, 25131, 12), dtype=uint8) type(y_true): <class 'tensorflow.python.framework.ops.Tensor'> y_pred: Tensor("sequential_26/time_distributed_26/Reshape_1:0", shape=(1, 25131, 12), dtype=float32) type(y_pred): <class 'tensorflow.python.framework.ops.Tensor'>
模型架构与数据维度
我的模型结构如下:
callbacks = [EarlyStopping(monitor='val_loss', patience=25)] model = Sequential() model.add(Masking(mask_value = 0.)) model.add(Bidirectional(LSTM(64, return_sequences=True, activation = "tanh"))) model.add(Dropout(0.3)) model.add(TimeDistributed(Dense(12, activation='softmax'))) adam = adam_v2.Adam(learning_rate=0.002) model.compile(optimizer=adam, loss=score, metrics=['accuracy']) history = model.fit(X_train, y_train, epochs=150, batch_size=1, shuffle=False, validation_data=(X_test, y_test), verbose=2, callbacks=[callbacks])
模型输入输出的维度如下,共12个分类:
print(f'{X_train.shape} {X_test.shape} {y_train.shape} {y_test.shape}') # 输出:(73, 25131, 29) (25, 23879, 29) (73, 25131, 12) (25, 23879, 12)
A为12×12的惩罚矩阵,维度与多分类任务的分类数对应:
该模型为岩性识别竞赛搭建。
解决方案
要实现适配TensorFlow张量、支持3维时序输入的自定义损失,按以下步骤调整即可:
- 首先把惩罚矩阵A转换为TensorFlow常量,保证可以在计算图中正常调用:
import tensorflow as tf # 假设A是你已经定义好的12*12 numpy数组 A_tensor = tf.constant(A, dtype=tf.float32)
- 重写自定义损失函数,替换numpy算子为TensorFlow原生算子,同时适配3维输入和Masking逻辑:
def custom_loss(y_true, y_pred): # 把独热编码的y_true转成类别索引,shape从(batch, 时序长度, 12)转为(batch, 时序长度) y_true_idx = tf.argmax(y_true, axis=-1, output_type=tf.int32) y_pred_idx = tf.argmax(y_pred, axis=-1, output_type=tf.int32) # 拼接索引得到每个样本点对应A矩阵的坐标 indices = tf.stack([y_true_idx, y_pred_idx], axis=-1) # 取出对应惩罚值,取负数和原指标逻辑对齐 penalty = -tf.gather_nd(A_tensor, indices) # 匹配Masking层逻辑,过滤全0的填充位置 mask = tf.cast(tf.reduce_any(y_true != 0, axis=-1), dtype=tf.float32) # 仅统计非填充位置的平均损失 loss = tf.reduce_sum(penalty * mask) / tf.reduce_sum(mask) return loss
- 替换模型编译时的损失函数即可:
model.compile(optimizer=adam, loss=custom_loss, metrics=['accuracy'])
优化建议
上面的损失直接使用argmax取预测类别,但argmax是不可导算子,可能导致训练梯度中断,收敛不稳定。更推荐使用带惩罚权重的交叉熵实现,既可以保留梯度,又能贴合你的惩罚逻辑:
def weighted_categorical_crossentropy(y_true, y_pred): y_true_idx = tf.argmax(y_true, axis=-1, output_type=tf.int32) # 取出真实类别对应的惩罚权重行 weights = tf.gather(A_tensor, y_true_idx) # 计算带惩罚权重的交叉熵 cce = tf.keras.losses.categorical_crossentropy(y_true, y_pred) weighted_cce = cce * tf.reduce_sum(y_true * weights, axis=-1) # 同样加入Mask逻辑 mask = tf.cast(tf.reduce_any(y_true != 0, axis=-1), dtype=tf.float32) loss = tf.reduce_sum(weighted_cce * mask) / tf.reduce_sum(mask) return loss
该实现训练收敛性更好,最终的指标表现也会更贴合你要优化的惩罚得分。
内容的提问来源于stack exchange,提问作者Matheus Schaly
相关产品推荐
相关产品推荐

