基于LSTM的收益率曲线预测模型形状不匹配报错求助
T+20收益率曲线LSTM预测模型的形状不匹配问题排查与优化
问题背景
本人是金融从业者,正在搭建用于预测T+20收益率曲线(10维向量)的AI模型,使用过去20天的收益率曲线数据作为输入。已将原始1001行×10列数据转换为(961,20,10)的输入数据和(961,10)的目标数据。使用带stateful=True的LSTM模型训练后,执行预测时出现形状不匹配错误。
模型代码
from keras.layers import LSTM, Dense import tensorflow as tf batch_size = 1 epochs = 2 timesteps = 20 inputs_1_mae = tf.keras.layers.Input(batch_shape=(batch_size, timesteps, 10)) lstm_1_mae = tf.keras.layers.LSTM(10, stateful = True, return_sequences = True)(inputs_1_mae) output_1_mae = tf.keras.layers.Dense(units = 10)(lstm_1_mae) regressor_mae = tf.keras.Model(inputs= inputs_1_mae ,outputs = output_1_mae) regressor_mae.compile (optimizer = "adam", loss = "mae") regressor_mae.summary() regressor_mae.fit(final_x_array, final_y_array, batch_size = batch_size, epochs=epochs)
模型结构输出
Model: "model_1" _________________________________________________________________ Layer (type) Output Shape Param # ================================================================= input_2 (InputLayer) [(1, 20, 10)] 0 lstm_2 (LSTM) (1, 20, 10) 840 dense_1 (Dense) (1, 20, 10) 110 ================================================================= Total params: 950 Trainable params: 950 Non-trainable params: 0
预测时的错误信息
InvalidArgumentError Traceback (most recent call last) ~\AppData\Local\Temp/ipykernel_240372/2358700359.py in <module> ----> 1 prediction = regressor_mae.predict(final_x_array) C:\ProgramData\Anaconda3\lib\site-packages\keras\utils\traceback_utils.py in error_handler(*args, **kwargs) 68 # To get the full stack trace, call: 69 # `tf.debugging.disable_traceback_filtering()` ---> 70 raise e.with_traceback(filtered_tb) from None 71 finally: 72 del filtered_tb C:\ProgramData\Anaconda3\lib\site-packages\tensorflow\python\eager\execute.py in quick_execute(op_name, num_outputs, inputs, attrs, ctx, name) 50 try: 51 ctx.ensure_initialized() ---> 52 tensors = pywrap_tfe.TFE_Py_Execute(ctx._handle, device_name, op_name, 53 inputs, attrs, num_outputs) 54 except core._NotOkStatusException as e: InvalidArgumentError: Graph execution error: Specified a list with shape [1,10] from a tensor with shape [32,10] [[{{node TensorArrayUnstack/TensorListFromTensor}}]] [[model_1/lstm_2/PartitionedCall]] [Op:__inference_predict_function_20772]
问题排查与解决方案
1. 形状不匹配的核心原因
- 输入输出形状不匹配:模型定义时输入
batch_shape=(1,20,10),且LSTM设置return_sequences=True,导致Dense层输出形状为(1,20,10),但训练目标数据是(961,10)——训练时batch_size=1会触发数据广播,但预测时predict方法默认用batch_size=32(TensorFlow默认值),和模型固定的batch_size=1冲突,触发形状错误。 - stateful LSTM的约束:带
stateful=True的LSTM要求训练和预测时的batch_size严格一致,且输入必须是固定batch大小的张量。
2. 即时修复方案
方案一:调整模型输出匹配目标形状
将LSTM的return_sequences=False,让LSTM只输出最后一个时间步的结果,Dense层输出形状变为(1,10),和目标数据(961,10)匹配:
inputs_1_mae = tf.keras.layers.Input(batch_shape=(batch_size, timesteps, 10)) # 修改return_sequences为False lstm_1_mae = tf.keras.layers.LSTM(10, stateful = True, return_sequences = False)(inputs_1_mae) output_1_mae = tf.keras.layers.Dense(units = 10)(lstm_1_mae) regressor_mae = tf.keras.Model(inputs= inputs_1_mae ,outputs = output_1_mae) regressor_mae.compile (optimizer = "adam", loss = "mae")
方案二:预测时指定batch_size=1
如果必须保留return_sequences=True(比如需要输出每个时间步的预测),则预测时明确指定batch_size=1,和模型定义一致:
prediction = regressor_mae.predict(final_x_array, batch_size=batch_size)
注意:此时需要将目标数据调整为(961,20,10),否则训练阶段也会存在潜在的形状不匹配问题。
3. 模型架构优化建议
(1)合理使用stateful LSTM
- stateful LSTM适合序列依赖极强的场景(比如时间序列连续预测),但需严格管理状态:训练结束后需调用
model.reset_states()重置状态,避免预测时携带训练的状态残留。 - 如果不需要跨batch保留状态,建议使用
stateful=False,模型输入无需固定batch_size,灵活性更高。
(2)提升网络容量
当前LSTM仅用10个单元,对于10维收益率曲线预测,容量可能不足。建议增加LSTM单元数(如32、64),并考虑堆叠多层LSTM:
inputs_1_mae = tf.keras.layers.Input(shape=(timesteps, 10)) # 不固定batch_size,stateful=False lstm_1 = tf.keras.layers.LSTM(64, return_sequences=True)(inputs_1_mae) lstm_2 = tf.keras.layers.LSTM(32, return_sequences=False)(lstm_1) output_1_mae = tf.keras.layers.Dense(10)(lstm_2)
(3)添加正则化与归一化
- 金融数据波动大,建议在输入层后添加
LayerNormalization或BatchNormalization稳定训练过程。 - 添加
Dropout层防止过拟合:
inputs_1_mae = tf.keras.layers.Input(shape=(timesteps, 10)) norm = tf.keras.layers.LayerNormalization()(inputs_1_mae) lstm_1 = tf.keras.layers.LSTM(64, return_sequences=True, dropout=0.2)(norm) lstm_2 = tf.keras.layers.LSTM(32, return_sequences=False, dropout=0.2)(lstm_1) output_1_mae = tf.keras.layers.Dense(10)(lstm_2)
(4)选择更适合的损失函数
除MAE外,金融场景下可使用Huber损失(结合MAE和MSE的优点,对异常值更鲁棒):
regressor_mae.compile(optimizer="adam", loss=tf.keras.losses.Huber())
内容的提问来源于stack exchange,提问作者Maagalam HARSHA VARDHAN
相关产品推荐
相关产品推荐

