股价预测:LSTM模型表现逊于常规回归模型,求技术指导
股价预测LSTM模型优化方案
问题背景
毕业论文聚焦股价预测,采用LSTM结合多技术指标开展研究,但当前常规回归模型表现优于LSTM,怀疑超参数或模型结构存在优化空间。
数据集信息
- 时间范围:2008年至今的日度数据
- 目标变量:每日收盘价
- 特征列表:
Open High Low Close Volume BOLL_Middle BOLL_Upper BOLL_Lower MACD CCI ATR EMA_20 MA5 MA10 MTM6 MTM12 ROC SMI WVAD USDX
当前LSTM模型代码
# Reshape the input data to be [samples, timesteps, features] train_features_reshaped = np.reshape(train_features.values, (train_features.shape[0], 1, train_features.shape[1])) val_features_reshaped = np.reshape(val_features.values, (val_features.shape[0], 1, val_features.shape[1])) test_features_reshaped = np.reshape(test_features.values, (test_features.shape[0], 1, test_features.shape[1])) # Determine the number of features in the input data num_columns = train_features.shape[1] # Define your custom learning rate learning_rate = 0.01 # Initialize the model model = Sequential() # First LSTM layer with Dropout regularization model.add(LSTM(units=100, return_sequences=True, input_shape=(1, num_columns))) model.add(Dropout(0.2)) # Second LSTM layer with Dropout regularization model.add(LSTM(units=100, return_sequences=True)) model.add(Dropout(0.2)) # Third LSTM layer with Dropout regularization model.add(LSTM(units=100)) model.add(Dropout(0.2)) # Fully Connected layer (Dense Layer) model.add(Dense(units=64, activation='relu')) # Output layer model.add(Dense(units=1)) # Compile the model with the custom learning rate optimizer = Adam(lr=learning_rate) model.compile(optimizer=optimizer, loss='mse') # Define the ModelCheckpoint callback to save the best model based on validation loss checkpoint_callback = ModelCheckpoint('best_model.h5', monitor='val_loss', save_best_only=True, mode='min', verbose=1) # Train the model and save the training history history = model.fit(train_features_reshaped, train_target, epochs=35, batch_size=64, validation_data=(val_features_reshaped, val_target), callbacks=[checkpoint_callback]) # Load the best model based on validation loss best_model = load_model('best_model.h5') # Evaluate the best model on the test set loss = best_model.evaluate(test_features_reshaped, test_target) # Make predictions using the best model predictions = best_model.predict(test_features_reshaped)
核心优化建议
1. 修复时间步长设置(关键问题)
当前代码timesteps=1,相当于只用单天特征预测次日收盘价,完全浪费了LSTM处理序列依赖的能力。股价是强时间序列,必须引入过去N天的历史数据:
- 尝试设置时间步长为7(一周)、14(两周)或30(一月),示例重构代码:
timesteps = 14 def create_sequences(features, target, timesteps): X, y = [], [] for i in range(timesteps, len(features)): X.append(features[i-timesteps:i]) y.append(target[i]) return np.array(X), np.array(y) train_features_reshaped, train_target = create_sequences(train_features.values, train_target.values, timesteps) val_features_reshaped, val_target = create_sequences(val_features.values, val_target.values, timesteps) test_features_reshaped, test_target = create_sequences(test_features.values, test_target.values, timesteps) - 时间步长不宜过大,否则会引入冗余噪声,可通过验证集损失筛选最优值。
2. 超参数精细化调整
学习率优化
当前learning_rate=0.01过高,Adam默认学习率为0.001,过高会导致模型训练震荡,难以收敛到最优解:
- 优先尝试0.001、0.0005、0.0001这几个常用值;
- 加入学习率自动调整回调:
from tensorflow.keras.callbacks import ReduceLROnPlateau lr_scheduler = ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=5, min_lr=1e-5) # 训练时加入该回调 history = model.fit(..., callbacks=[checkpoint_callback, lr_scheduler])
LSTM单元与层数调整
当前3层100单元的结构可能冗余或过拟合:
- 先尝试减少层数(如1-2层),观察验证集表现;
- 单元数量测试32、64、128三个档位,配合验证集损失筛选最优值。
Dropout策略优化
固定0.2的Dropout比例可调整:
- 尝试0.1-0.3区间的不同值;
- 改用
recurrent_dropout针对LSTM循环层做正则化(注意会增加训练时间):model.add(LSTM(units=64, return_sequences=True, input_shape=(timesteps, num_columns), recurrent_dropout=0.2))
3. 数据预处理强化
特征标准化
股价、成交量等特征数值范围差异极大,LSTM对数据尺度敏感,必须做标准化:
- 用
StandardScaler或MinMaxScaler对特征和目标变量分别处理,注意仅用训练集拟合scaler,验证/测试集复用训练集的scaler,避免数据泄露; - 预测后需将标准化的结果反转换回原始尺度,再计算评估指标。
特征筛选
20个特征存在冗余(如MA5与MA10相关性极高):
- 计算特征与目标变量的相关性,筛选Top10-15的特征;
- 用PCA或互信息法降低特征维度,减少噪声干扰。
4. 模型结构升级
加入注意力机制
让LSTM自动关注对预测更重要的时间步:
from tensorflow.keras.layers import Attention, Permute, RepeatVector, Multiply, GlobalAveragePooling1D, Flatten, Activation # 假设最后一层LSTM是模型的倒数第二层 lstm_output = model.layers[-2].output # 构建注意力权重 attention = Dense(1, activation='tanh')(lstm_output) attention = Flatten()(attention) attention = Activation('softmax')(attention) attention = RepeatVector(num_columns)(attention) attention = Permute([2, 1])(attention) # 加权输出 weighted_output = Multiply()([lstm_output, attention]) weighted_output = GlobalAveragePooling1D()(weighted_output) # 连接全连接层与输出层 x = Dense(64, activation='relu')(weighted_output) output = Dense(1)(x) # 重新定义模型 model = Model(inputs=model.input, outputs=output)
5. 训练策略优化
增加训练轮数+早停
当前35轮训练可能未收敛,配合早停防止过拟合:
from tensorflow.keras.callbacks import EarlyStopping early_stopping = EarlyStopping(monitor='val_loss', patience=10, restore_best_weights=True) history = model.fit(..., callbacks=[checkpoint_callback, lr_scheduler, early_stopping])
- 初始训练轮数可设为100-200,由早停机制自动终止。
调整Batch Size
当前batch_size=64,可尝试32、128,较小的batch_size可能提升泛化性,但训练速度会变慢。
6. 基准对比规范
- 确保常规回归模型与LSTM使用完全相同的数据集划分、预处理流程,避免因数据差异导致结果偏差;
- 除MSE外,补充MAE、RMSE、R²等评估指标,LSTM可能在趋势预测上表现更好,即使MSE略高于回归模型。
内容的提问来源于stack exchange,提问作者Samir Benchouk
相关产品推荐
相关产品推荐

