Keras Softmax层输出概率和不为1的技术求助
四类多分类任务Softmax输出概率和不为1的问题排查与解决
核心原因分析
- 混合精度数值溢出:
mixed_float16的动态范围远小于float32,当softmax的输入特征中有极大值时,计算exp会触发溢出,导致部分数值被截断为0或1,最终概率和偏离1。 - 模型结构梯度问题:在预训练LM输出后叠加双向LSTM,增加了模型深度,容易引发梯度消失,导致分类头权重更新异常,softmax无法正常完成归一化。
具体解决方案
1. 修复混合精度的数值稳定性
不要全局强制所有层使用float16,尤其是最后一层softmax,需强制用float32计算避免溢出:
# 修改分类头的Dense层 output = tf.keras.layers.Dense(4, activation='softmax', dtype=tf.float32 # 强制使用float32计算 )(out)
同时,使用混合精度的标准流程,配合LossScaleOptimizer优化器,从根源避免数值下溢/溢出:
optimizer = tf.keras.optimizers.Adam() # 包装优化器,启用损失缩放 optimizer = tf.keras.mixed_precision.LossScaleOptimizer(optimizer)
2. 简化模型结构,避免梯度消失
预训练LM的输出已经包含足够的语义信息,无需额外叠加LSTM,直接池化后拼接即可:
# 替换原有的LSTM模块 # Text input 1 inputs_descr = keras.Input(shape=(seq_length_descr,), dtype=tf.int32, name='input_1') out_descr = pretrained_lm.layers[1](inputs_descr) out_descr = tf.keras.layers.GlobalMaxPool1D()(out_descr) # 直接池化 # Text Input 2 inputs_tw = keras.Input(shape=(seq_length_tw,), dtype=tf.int32, name='input_2') out_tw = pretrained_lm.layers[1](inputs_tw) out_tw = tf.keras.layers.GlobalMaxPool1D()(out_tw)
若坚持保留LSTM,需在LSTM前添加层归一化,稳定数值分布:
out_descr = pretrained_lm.layers[1](inputs_descr) out_descr = tf.keras.layers.LayerNormalization()(out_descr) # 添加层归一化 out_descr = Bidirectional(LSTM(50, return_sequences=True, activation='tanh', dropout=0.2))(out_descr) out_descr = tf.keras.layers.GlobalMaxPool1D()(out_descr)
3. 推理阶段的临时修复
如果训练暂时无法调整,可以在推理时手动用float32重新计算softmax,得到正确的归一化结果:
import numpy as np import tensorflow as tf raw_preds = model.predict(inputs) # 转换为float32后重新计算softmax normalized_preds = tf.nn.softmax(raw_preds.astype(np.float32)).numpy() # 验证概率和 print(normalized_preds.sum(axis=1)) # 结果应接近1
内容的提问来源于stack exchange,提问作者Santiago Esteban
相关产品推荐
相关产品推荐

