Optuna优化PyTorch LSTM超参数使用nn.LSTMCell报错如何解决
问题根因
报错触发的核心原因有两个:
nn.Sequential的执行逻辑是将前一层的单个返回张量直接传入下一层作为输入,但nn.LSTMCell的前向传播要求输入为(当前步输入张量, 上一时刻隐藏状态元组(h,c)),返回值也是(更新后的h, 更新后的c)元组,直接放入Sequential会导致后续层接收到元组而非张量,调用.dim()方法时触发属性错误。- 代码存在笔误:Optuna传入的参数对象名为
trial,代码中误写为trail,运行时会触发参数不存在的二级错误。
原错误实现代码如下:
def build_model_custom(trail): # Suggest the number of layers of neural network model n_layers = trail.suggest_int("n_layers", 1, 3) layers = [] in_features = 20 for i in range(n_layers): # Suggest the number of units in each layer out_features = trail.suggest_int("n_units_l{}".format(i), 4, 18) layers.append(nn.LSTMCell(in_features, out_features)) in_features = out_features layers.append(nn.Linear(in_features, 2)) return nn.Sequential(*layers)
修复方案
你可以根据需求选择两种修复方式,两种方式都兼容Optuna的超参数采样逻辑:
方案1:保留nn.LSTMCell,自定义模型类管理前向逻辑
nn.LSTMCell是单时间步的LSTM计算单元,需要手动维护每层的隐藏状态、逐时间步循环计算,不能直接用Sequential拼接,自定义模型实现如下:
import torch import torch.nn as nn class LSTMModel(nn.Module): def __init__(self, trial, input_dim=20, output_dim=2): super().__init__() # 超参数采样逻辑保留,修正trial拼写 self.n_layers = trial.suggest_int("n_layers", 1, 3) self.lstm_cells = nn.ModuleList() in_feat = input_dim for i in range(self.n_layers): out_feat = trial.suggest_int(f"n_units_l{i}", 4, 18) self.lstm_cells.append(nn.LSTMCell(in_feat, out_feat)) in_feat = out_feat self.fc = nn.Linear(in_feat, output_dim) def forward(self, x): # 输入x形状: (batch_size, 序列长度, 输入维度20) batch_size = x.shape[0] device = x.device # 初始化每层的h、c状态 state_list = [] for cell in self.lstm_cells: h = torch.zeros(batch_size, cell.hidden_size, device=device) c = torch.zeros(batch_size, cell.hidden_size, device=device) state_list.append((h, c)) # 逐时间步计算 for t in range(x.shape[1]): cur_input = x[:, t, :] for idx, cell in enumerate(self.lstm_cells): h_prev, c_prev = state_list[idx] h_new, c_new = cell(cur_input, (h_prev, c_prev)) state_list[idx] = (h_new, c_new) cur_input = h_new # 取最后一层最后时刻的h传入全连接层 return self.fc(state_list[-1][0]) # Optuna调用入口 def build_model_custom(trial): return LSTMModel(trial)
方案2:用nn.LSTM替换nn.LSTMCell,适配Sequential逻辑
如果不需要自定义单时间步的计算逻辑,直接使用多层LSTM模块nn.LSTM即可,它支持直接输入整个序列,只需加一层包装取出需要的输出张量,就可以放入Sequential中:
import torch.nn as nn def build_model_custom(trial): n_layers = trial.suggest_int("n_layers", 1, 3) hidden_dim = trial.suggest_int("hidden_dim", 4, 18) # 包装LSTM层,适配Sequential的单张量输入输出要求 class LSTMWrap(nn.Module): def __init__(self): super().__init__() self.lstm = nn.LSTM( input_size=20, hidden_size=hidden_dim, num_layers=n_layers, batch_first=True ) def forward(self, x): # lstm返回(序列所有时刻输出, (最终h, 最终c)),取最后一个时刻的输出 seq_out, _ = self.lstm(x) return seq_out[:, -1, :] return nn.Sequential( LSTMWrap(), nn.Linear(hidden_dim, 2) )
内容的提问来源于stack exchange,提问作者Joe_Joe
相关产品推荐
相关产品推荐

