You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch LSTM模型num_layers参数用法及输出维度选择问题

问题根源

PyTorch内置nn.LSTM返回的最终隐藏状态张量维度不受batch_first参数影响,固定为(num_layers * num_directions, batch_size, hidden_dim):

  • 单层LSTM场景下num_layers=1,该张量第一维长度为1,直接送入全连接层时维度刚好匹配,所以之前运行结果符合预期。
  • 双层LSTM场景下num_layers=2,该张量第一维长度为2,第一维的两个元素分别对应第一层、第二层LSTM在序列最后一个时间步的输出隐藏状态,直接送入全连接层就会得到你遇到的[2,8]维度结果,和预期不符。
有效输出选取逻辑

多层LSTM的计算逻辑是下层每个时间步的输出作为上层对应时间步的输入,只有最顶层LSTM在最后一个时间步的隐藏状态,是经过全部网络层提取后的最终序列特征,也就是需要送入后续全连接层的有效内容,通过h_out[-1, :, :]即可索引到,该张量维度为(batch_size, hidden_dim),完全匹配全连接层的输入维度要求。

修正后的代码

你现有代码只需要修改forward方法中隐藏状态的提取逻辑即可,另外可以移除冗余的Variable包装和多余的cuda调用:

class LSTM(nn.Module):
    def __init__(self, input_dim, hidden_dim, num_layers, output_dim):
        super(LSTM, self).__init__()
        self.input_dim = input_dim
        self.output_dim = output_dim
        self.hidden_dim = hidden_dim
        self.num_layers = num_layers

        self.lstm = nn.LSTM(input_size=input_dim, hidden_size=hidden_dim,
                            num_layers=num_layers, batch_first=True)
        self.fc = nn.Linear(hidden_dim, output_dim)

    def forward(self, x):
        h_0 = torch.zeros(
            self.num_layers, x.size(0), self.hidden_dim).cuda()
        c_0 = torch.zeros(
            self.num_layers, x.size(0), self.hidden_dim).cuda()
        
        # 前向传播计算LSTM输出
        ula, (h_out, _) = self.lstm(x, (h_0, c_0))
        # 提取最后一层的最终隐藏状态
        h_out = h_out[-1, :, :]
        out = self.fc(h_out)
        return out

补充:如果你需要用每个时间步的输出做序列标注类任务,可以直接用LSTM返回的第一个返回值ula,当设置batch_first=True时它的维度是(batch_size, seq_len, hidden_dim),对应每个时间步最顶层LSTM的输出。

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 13:03:21