PyTorch训练LSTM时遭遇梯度消失问题求助
PyTorch LSTM股票价格预测模型无法拟合:损失爆炸、R2为负的问题排查
我正在用PyTorch训练一个简单LSTM神经网络做股票价格预测,但模型完全无法拟合——损失爆炸、R2值为负,训练过程毫无改善。代码里肯定有致命错误,试了多种方法都没解决。
以下是我的代码:
class LSTMModel(nn.Module): def __init__(self, features): super(LSTMModel, self).__init__() self.lstm1 = nn.LSTM(input_size=features, hidden_size=16, batch_first=True) self.dense2 = nn.Linear(16, 1) self._init_weights() def forward(self, x): x, _ = self.lstm1(x) # x, _ = self.lstm2(x) # Flatten the output for Dense layer input x = x[:, -1, :] # x = self.dense1(x) x = self.dense2(x) return x def _init_weights(self): for name, param in self.named_parameters(): if 'weight' in name: nn.init.xavier_uniform_(param) elif 'bias' in name: nn.init.zeros_(param) # Initialize the model model = LSTMModel(len(feature_cols)) criterion = nn.MSELoss() optimizer = optim.Adam(model.parameters(), lr=0.01) scheduler = optim.lr_scheduler.ExponentialLR(optimizer, gamma=0.95) def train_model(num_epochs): for epoch in range(num_epochs): model.train() total_loss = 0 for data, target in train_loader: optimizer.zero_grad() output = model(data) loss = criterion(output.reshape(len(output), ), target) loss.backward() torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) optimizer.step() total_loss += loss.item() scheduler.step() model.eval() val_loss = 0 pred = [] with torch.no_grad(): for data, target in test_loader: output = model(data) # print(data, output.reshape(len(output), ), target) val_loss += criterion(output.reshape(len(output), ), target).item() pred += list(output.reshape(len(output), )) val_loss /= len(test_loader) r2 = r2_score(test_y, pred) print(f'Epoch {epoch + 1}, Train Loss: {total_loss / len(train_loader)}, Val Loss: {val_loss}, val r2: {r2}')
我已尝试的方法:
- 梯度裁剪(代码中已实现),无效
- 修改batch size,无效
- 查看网络权重,发现LSTM层权重均接近0,而全连接层权重正常,疑似梯度消失问题
- 自定义权重初始化(代码中已实现),无效
- 修改模型超参数(包括隐藏层数量、学习率、hidden_size等),无效
- 修改输入特征数量,无效
- 调整时间序列数据的滑动窗口大小,无效
备注:
- 输入数据特征已使用MinMaxScaler进行归一化
- 数据集包含约4000条观测数据
内容的提问来源于stack exchange,提问作者王一诺
相关产品推荐
相关产品推荐

