PyTorch LSTM模型num_layers参数用法及输出维度选择问题
问题根源
PyTorch内置nn.LSTM返回的最终隐藏状态张量维度不受batch_first参数影响,固定为(num_layers * num_directions, batch_size, hidden_dim):
- 单层LSTM场景下
num_layers=1,该张量第一维长度为1,直接送入全连接层时维度刚好匹配,所以之前运行结果符合预期。 - 双层LSTM场景下
num_layers=2,该张量第一维长度为2,第一维的两个元素分别对应第一层、第二层LSTM在序列最后一个时间步的输出隐藏状态,直接送入全连接层就会得到你遇到的[2,8]维度结果,和预期不符。
有效输出选取逻辑
多层LSTM的计算逻辑是下层每个时间步的输出作为上层对应时间步的输入,只有最顶层LSTM在最后一个时间步的隐藏状态,是经过全部网络层提取后的最终序列特征,也就是需要送入后续全连接层的有效内容,通过h_out[-1, :, :]即可索引到,该张量维度为(batch_size, hidden_dim),完全匹配全连接层的输入维度要求。
修正后的代码
你现有代码只需要修改forward方法中隐藏状态的提取逻辑即可,另外可以移除冗余的Variable包装和多余的cuda调用:
class LSTM(nn.Module): def __init__(self, input_dim, hidden_dim, num_layers, output_dim): super(LSTM, self).__init__() self.input_dim = input_dim self.output_dim = output_dim self.hidden_dim = hidden_dim self.num_layers = num_layers self.lstm = nn.LSTM(input_size=input_dim, hidden_size=hidden_dim, num_layers=num_layers, batch_first=True) self.fc = nn.Linear(hidden_dim, output_dim) def forward(self, x): h_0 = torch.zeros( self.num_layers, x.size(0), self.hidden_dim).cuda() c_0 = torch.zeros( self.num_layers, x.size(0), self.hidden_dim).cuda() # 前向传播计算LSTM输出 ula, (h_out, _) = self.lstm(x, (h_0, c_0)) # 提取最后一层的最终隐藏状态 h_out = h_out[-1, :, :] out = self.fc(h_out) return out
补充:如果你需要用每个时间步的输出做序列标注类任务,可以直接用LSTM返回的第一个返回值
ula,当设置batch_first=True时它的维度是(batch_size, seq_len, hidden_dim),对应每个时间步最顶层LSTM的输出。
内容的提问来源于stack exchange,提问作者Joe
相关产品推荐
相关产品推荐

