基于Keras的Transformer模型测试损失无变化问题求助
股票价格预测Transformer模型欠拟合问题求助
基于Keras官方时序Transformer分类教程搭建了股票价格预测模型,遇到严重欠拟合问题:测试损失数值极大,训练各轮次间损失几乎无变化,模型输出全为相同值。但用相同数据搭建同结构的LSTM模型时,损失能持续下降且预测效果良好。尝试调整学习率、堆叠更多transformer_encoder_block模块均无改善,求解决建议。
模型代码
def transformer_encoder_block(inputs, head_size, num_heads, filters, dropout=0): # Normalization and Attention x = layers.LayerNormalization(epsilon=1e-6)(inputs) x = layers.MultiHeadAttention( key_dim=head_size, num_heads=num_heads, dropout=dropout )(x, x) x = layers.Dropout(dropout)(x) res = x + inputs # Feed Forward Part x = layers.LayerNormalization(epsilon=1e-6)(res) x = layers.Conv1D(filters=filters, kernel_size=1, activation="relu")(x) x = layers.Dropout(dropout)(x) x = layers.Conv1D(filters=inputs.shape[-1], kernel_size=1)(x) return x + res data = ... input = np.array( keras.preprocessing.sequence.pad_sequences(data["input"], padding="pre", dtype="float32")) output = np.array( keras.preprocessing.sequence.pad_sequences(data["output"], padding="pre", dtype="float32")) # Input shape: (723, 36, 22) # Output shape: (723, 36, 1) # Train data train_features = input[100:] train_labels = output[100:] train_labels = tf.keras.utils.to_categorical(train_labels, num_classes=3) # Test data test_features = input[:100] test_labels = output[:100] test_labels = tf.keras.utils.to_categorical(test_labels, num_classes=3) inputs = keras.Input(shape=(None,22), dtype="float32", name="inputs") # Ignore padding in inputs x = layers.Masking(mask_value=0)(inputs) x = transformer_encoder_block(x, head_size=64, num_heads=16, filters=3, dropout=0.2) # Multiclass = Softmax (decrease, no change, increase) outputs = layers.TimeDistributed(layers.Dense(3, activation="softmax", name="outputs"))(x) # Create model model = keras.Model(inputs=inputs, outputs=outputs) # Compile model model.compile(loss="categorical_crossentropy", optimizer=(tf.keras.optimizers.Adam(learning_rate=0.005)), metrics=['accuracy']) # Train model history = model.fit(train_features, train_labels, epochs=10, batch_size=32) # Evaluate on the test data test_loss = model.evaluate(test_features, test_labels, verbose=0) print("Test loss:", test_loss) out = model.predict(test_features)
数据与训练情况
填充后输入形状为(723, 36, 22),输出形状为(723, 36, 1),转换为独热编码后对应3个输出类别(下跌、持平、上涨)。
10轮训练输出示例(训练更多轮次无改善):
Epoch 1/10 20/20 [==============================] - 2s 62ms/step - loss: 10.7436 - accuracy: 0.3335 Epoch 2/10 20/20 [==============================] - 1s 62ms/step - loss: 10.7083 - accuracy: 0.3354 Epoch 3/10 20/20 [==============================] - 1s 60ms/step - loss: 10.6555 - accuracy: 0.3392 Epoch 4/10 20/20 [==============================] - 1s 62ms/step - loss: 10.7846 - accuracy: 0.3306 Epoch 5/10 20/20 [==============================] - 1s 60ms/step - loss: 10.7600 - accuracy: 0.3322 Epoch 6/10 20/20 [==============================] - 1s 59ms/step - loss: 10.7074 - accuracy: 0.3358 Epoch 7/10 20/20 [==============================] - 1s 59ms/step - loss: 10.6569 - accuracy: 0.3385 Epoch 8/10 20/20 [==============================] - 1s 60ms/step - loss: 10.7767 - accuracy: 0.3314 Epoch 9/10 20/20 [==============================] - 1s 61ms/step - loss: 10.7346 - accuracy: 0.3341 Epoch 10/10 20/20 [==============================] - 1s 62ms/step - loss: 10.7093 - accuracy: 0.3354 Test loss: [10.073813438415527, 0.375] 4/4 [==============================] - 0s 22ms/step
解决建议
- 特征归一化:Transformer对数据尺度极度敏感,LSTM相对鲁棒。先对输入特征做标准化处理(如
StandardScaler),消除不同特征间的数值差异,避免注意力机制被大尺度特征主导。 - 修正注意力模块参数:当前
head_size=64、num_heads=16,总维度为64*16=1024,远大于输入特征维度22,导致注意力模块无法有效捕捉特征关联。建议调整参数,让head_size * num_heads接近输入特征维度,比如head_size=11、num_heads=2(总维度22),或head_size=8、num_heads=3(总维度24)。 - 调整Feed Forward层容量:当前
Conv1D(filters=3)过小,特征压缩过度,无法有效传递信息。将filters调至64、128等更大数值,增强FFN的特征处理能力。 - 添加位置编码:Transformer本身没有时序感知能力,必须手动注入位置信息。在
Masking层后添加位置编码,示例代码如下:class PositionalEncoding(layers.Layer): def __init__(self, position, d_model): super().__init__() self.pos_encoding = self.positional_encoding(position, d_model) def get_angles(self, position, i, d_model): angles = 1 / tf.pow(10000, (2 * (i // 2)) / tf.cast(d_model, tf.float32)) return position * angles def positional_encoding(self, position, d_model): angle_rads = self.get_angles( position=tf.range(position, dtype=tf.float32)[:, tf.newaxis], i=tf.range(d_model, dtype=tf.float32)[tf.newaxis, :], d_model=d_model ) sines = tf.math.sin(angle_rads[:, 0::2]) cosines = tf.math.cos(angle_rads[:, 1::2]) pos_encoding = tf.concat([sines, cosines], axis=-1) pos_encoding = pos_encoding[tf.newaxis, ...] return tf.cast(pos_encoding, tf.float32) # 在Masking层后添加 max_seq_len = 36 # 你的序列长度 pos_encoder = PositionalEncoding(max_seq_len, 22) x = x + pos_encoder(x) - 优化学习率设置:当前学习率0.005过高,容易导致训练震荡。尝试降至0.001或0.0001,同时配合
ReduceLROnPlateau回调函数,根据验证损失动态调整学习率。 - 检查标签处理逻辑:确认
to_categorical转换正确,原输出的数值需是0、1、2的整数类别,避免因标签格式错误导致损失计算异常。 - 合理堆叠编码器模块:之前堆叠无效可能是参数设置问题,调整上述参数后,尝试堆叠2-3个
transformer_encoder_block,提升模型容量。
内容的提问来源于stack exchange,提问作者Twisted Tea
相关产品推荐
相关产品推荐

