You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

股价预测:LSTM模型表现逊于常规回归模型,求技术指导

股价预测LSTM模型优化方案

问题背景

毕业论文聚焦股价预测,采用LSTM结合多技术指标开展研究,但当前常规回归模型表现优于LSTM,怀疑超参数或模型结构存在优化空间。

数据集信息

  • 时间范围:2008年至今的日度数据
  • 目标变量:每日收盘价
  • 特征列表:Open High Low Close Volume BOLL_Middle BOLL_Upper BOLL_Lower MACD CCI ATR EMA_20 MA5 MA10 MTM6 MTM12 ROC SMI WVAD USDX

当前LSTM模型代码

# Reshape the input data to be [samples, timesteps, features]
train_features_reshaped = np.reshape(train_features.values, (train_features.shape[0], 1, train_features.shape[1]))
val_features_reshaped = np.reshape(val_features.values, (val_features.shape[0], 1, val_features.shape[1]))
test_features_reshaped = np.reshape(test_features.values, (test_features.shape[0], 1, test_features.shape[1]))

# Determine the number of features in the input data
num_columns = train_features.shape[1]

# Define your custom learning rate
learning_rate = 0.01

# Initialize the model
model = Sequential()

# First LSTM layer with Dropout regularization
model.add(LSTM(units=100, return_sequences=True, input_shape=(1, num_columns)))
model.add(Dropout(0.2))

# Second LSTM layer with Dropout regularization
model.add(LSTM(units=100, return_sequences=True))
model.add(Dropout(0.2))

# Third LSTM layer with Dropout regularization
model.add(LSTM(units=100))
model.add(Dropout(0.2))

# Fully Connected layer (Dense Layer)
model.add(Dense(units=64, activation='relu'))

# Output layer
model.add(Dense(units=1))

# Compile the model with the custom learning rate
optimizer = Adam(lr=learning_rate)
model.compile(optimizer=optimizer, loss='mse')

# Define the ModelCheckpoint callback to save the best model based on validation loss
checkpoint_callback = ModelCheckpoint('best_model.h5', monitor='val_loss', save_best_only=True, mode='min', verbose=1)

# Train the model and save the training history
history = model.fit(train_features_reshaped, train_target, epochs=35, batch_size=64, 
                    validation_data=(val_features_reshaped, val_target),
                    callbacks=[checkpoint_callback])

# Load the best model based on validation loss
best_model = load_model('best_model.h5')

# Evaluate the best model on the test set
loss = best_model.evaluate(test_features_reshaped, test_target)

# Make predictions using the best model
predictions = best_model.predict(test_features_reshaped)

核心优化建议

1. 修复时间步长设置(关键问题)

当前代码timesteps=1,相当于只用单天特征预测次日收盘价,完全浪费了LSTM处理序列依赖的能力。股价是强时间序列,必须引入过去N天的历史数据:

  • 尝试设置时间步长为7(一周)、14(两周)或30(一月),示例重构代码:
    timesteps = 14
    def create_sequences(features, target, timesteps):
        X, y = [], []
        for i in range(timesteps, len(features)):
            X.append(features[i-timesteps:i])
            y.append(target[i])
        return np.array(X), np.array(y)
    
    train_features_reshaped, train_target = create_sequences(train_features.values, train_target.values, timesteps)
    val_features_reshaped, val_target = create_sequences(val_features.values, val_target.values, timesteps)
    test_features_reshaped, test_target = create_sequences(test_features.values, test_target.values, timesteps)
    
  • 时间步长不宜过大,否则会引入冗余噪声,可通过验证集损失筛选最优值。

2. 超参数精细化调整

学习率优化

当前learning_rate=0.01过高,Adam默认学习率为0.001,过高会导致模型训练震荡,难以收敛到最优解:

  • 优先尝试0.001、0.0005、0.0001这几个常用值;
  • 加入学习率自动调整回调:
    from tensorflow.keras.callbacks import ReduceLROnPlateau
    lr_scheduler = ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=5, min_lr=1e-5)
    # 训练时加入该回调
    history = model.fit(..., callbacks=[checkpoint_callback, lr_scheduler])
    

LSTM单元与层数调整

当前3层100单元的结构可能冗余或过拟合:

  • 先尝试减少层数(如1-2层),观察验证集表现;
  • 单元数量测试32、64、128三个档位,配合验证集损失筛选最优值。

Dropout策略优化

固定0.2的Dropout比例可调整:

  • 尝试0.1-0.3区间的不同值;
  • 改用recurrent_dropout针对LSTM循环层做正则化(注意会增加训练时间):
    model.add(LSTM(units=64, return_sequences=True, input_shape=(timesteps, num_columns), recurrent_dropout=0.2))
    

3. 数据预处理强化

特征标准化

股价、成交量等特征数值范围差异极大,LSTM对数据尺度敏感,必须做标准化:

  • 用StandardScaler或MinMaxScaler对特征和目标变量分别处理,注意仅用训练集拟合scaler,验证/测试集复用训练集的scaler,避免数据泄露;
  • 预测后需将标准化的结果反转换回原始尺度,再计算评估指标。

特征筛选

20个特征存在冗余(如MA5与MA10相关性极高):

  • 计算特征与目标变量的相关性,筛选Top10-15的特征;
  • 用PCA或互信息法降低特征维度,减少噪声干扰。

4. 模型结构升级

加入注意力机制

让LSTM自动关注对预测更重要的时间步:

from tensorflow.keras.layers import Attention, Permute, RepeatVector, Multiply, GlobalAveragePooling1D, Flatten, Activation

# 假设最后一层LSTM是模型的倒数第二层
lstm_output = model.layers[-2].output
# 构建注意力权重
attention = Dense(1, activation='tanh')(lstm_output)
attention = Flatten()(attention)
attention = Activation('softmax')(attention)
attention = RepeatVector(num_columns)(attention)
attention = Permute([2, 1])(attention)
# 加权输出
weighted_output = Multiply()([lstm_output, attention])
weighted_output = GlobalAveragePooling1D()(weighted_output)
# 连接全连接层与输出层
x = Dense(64, activation='relu')(weighted_output)
output = Dense(1)(x)
# 重新定义模型
model = Model(inputs=model.input, outputs=output)

5. 训练策略优化

增加训练轮数+早停

当前35轮训练可能未收敛,配合早停防止过拟合:

from tensorflow.keras.callbacks import EarlyStopping
early_stopping = EarlyStopping(monitor='val_loss', patience=10, restore_best_weights=True)
history = model.fit(..., callbacks=[checkpoint_callback, lr_scheduler, early_stopping])
  • 初始训练轮数可设为100-200,由早停机制自动终止。

调整Batch Size

当前batch_size=64,可尝试32、128,较小的batch_size可能提升泛化性,但训练速度会变慢。

6. 基准对比规范

  • 确保常规回归模型与LSTM使用完全相同的数据集划分、预处理流程,避免因数据差异导致结果偏差;
  • 除MSE外,补充MAE、RMSE、R²等评估指标,LSTM可能在趋势预测上表现更好,即使MSE略高于回归模型。

内容的提问来源于stack exchange,提问作者Samir Benchouk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 15:53:13