You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Optuna优化PyTorch LSTM超参数使用nn.LSTMCell报错如何解决

问题根因

报错触发的核心原因有两个:

  1. nn.Sequential的执行逻辑是将前一层的单个返回张量直接传入下一层作为输入,但nn.LSTMCell的前向传播要求输入为(当前步输入张量, 上一时刻隐藏状态元组(h,c)),返回值也是(更新后的h, 更新后的c)元组,直接放入Sequential会导致后续层接收到元组而非张量,调用.dim()方法时触发属性错误。
  2. 代码存在笔误:Optuna传入的参数对象名为trial,代码中误写为trail,运行时会触发参数不存在的二级错误。

原错误实现代码如下:

def build_model_custom(trail):
    
    # Suggest the number of layers of neural network model
    n_layers = trail.suggest_int("n_layers", 1, 3)
    layers = []

    in_features = 20
    
    for i in range(n_layers):
        
        # Suggest the number of units in each layer
        out_features = trail.suggest_int("n_units_l{}".format(i), 4, 18)
        
        layers.append(nn.LSTMCell(in_features, out_features))

        in_features = out_features
        
    layers.append(nn.Linear(in_features, 2))

    return nn.Sequential(*layers)
修复方案

你可以根据需求选择两种修复方式,两种方式都兼容Optuna的超参数采样逻辑:

方案1:保留nn.LSTMCell,自定义模型类管理前向逻辑

nn.LSTMCell是单时间步的LSTM计算单元,需要手动维护每层的隐藏状态、逐时间步循环计算,不能直接用Sequential拼接,自定义模型实现如下:

import torch
import torch.nn as nn

class LSTMModel(nn.Module):
    def __init__(self, trial, input_dim=20, output_dim=2):
        super().__init__()
        # 超参数采样逻辑保留,修正trial拼写
        self.n_layers = trial.suggest_int("n_layers", 1, 3)
        self.lstm_cells = nn.ModuleList()
        in_feat = input_dim

        for i in range(self.n_layers):
            out_feat = trial.suggest_int(f"n_units_l{i}", 4, 18)
            self.lstm_cells.append(nn.LSTMCell(in_feat, out_feat))
            in_feat = out_feat
        
        self.fc = nn.Linear(in_feat, output_dim)

    def forward(self, x):
        # 输入x形状: (batch_size, 序列长度, 输入维度20)
        batch_size = x.shape[0]
        device = x.device
        # 初始化每层的h、c状态
        state_list = []
        for cell in self.lstm_cells:
            h = torch.zeros(batch_size, cell.hidden_size, device=device)
            c = torch.zeros(batch_size, cell.hidden_size, device=device)
            state_list.append((h, c))
        
        # 逐时间步计算
        for t in range(x.shape[1]):
            cur_input = x[:, t, :]
            for idx, cell in enumerate(self.lstm_cells):
                h_prev, c_prev = state_list[idx]
                h_new, c_new = cell(cur_input, (h_prev, c_prev))
                state_list[idx] = (h_new, c_new)
                cur_input = h_new
        
        # 取最后一层最后时刻的h传入全连接层
        return self.fc(state_list[-1][0])

# Optuna调用入口
def build_model_custom(trial):
    return LSTMModel(trial)

方案2:用nn.LSTM替换nn.LSTMCell,适配Sequential逻辑

如果不需要自定义单时间步的计算逻辑,直接使用多层LSTM模块nn.LSTM即可,它支持直接输入整个序列,只需加一层包装取出需要的输出张量,就可以放入Sequential中:

import torch.nn as nn

def build_model_custom(trial):
    n_layers = trial.suggest_int("n_layers", 1, 3)
    hidden_dim = trial.suggest_int("hidden_dim", 4, 18)

    # 包装LSTM层,适配Sequential的单张量输入输出要求
    class LSTMWrap(nn.Module):
        def __init__(self):
            super().__init__()
            self.lstm = nn.LSTM(
                input_size=20,
                hidden_size=hidden_dim,
                num_layers=n_layers,
                batch_first=True
            )
        def forward(self, x):
            # lstm返回(序列所有时刻输出, (最终h, 最终c)),取最后一个时刻的输出
            seq_out, _ = self.lstm(x)
            return seq_out[:, -1, :]
    
    return nn.Sequential(
        LSTMWrap(),
        nn.Linear(hidden_dim, 2)
    )

内容的提问来源于stack exchange,提问作者Joe_Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 11:39:42