PyTorch复现Coursera TensorFlow LSTM时序预测模型问题咨询
PyTorch复现Coursera时序预测模型修正方案
你现有代码存在多处维度、层逻辑的错误,以下是逐点问题说明和可直接运行的正确实现。
现有代码核心问题
- 输入与卷积层维度逻辑颠倒:原TensorFlow模型处理的是单变量时序,输入形状为
(批次大小, 滑动窗口长度30, 特征数1),你误将批次维度32当成了卷积输入通道数,测试输入的形状构造完全错误。 - 卷积层padding配置错误:原模型用
padding="causal"(因果卷积,保证预测时不会泄露未来信息),你直接设padding=2会在序列两端补零,导致输出序列长度变长,且不符合因果卷积要求。 - LSTM层调用与参数逻辑错误:第一层LSTM对应TF的
return_sequences=True需要返回所有时间步输出,第二层LSTM对应TF默认配置只需要返回最后一个时间步的隐状态;你代码中调用层时误写为self.LSTM(和类名重名)、括号未闭合,且错误地展平了所有时间步特征,把批次维度和序列维度混在了一起。 - 全连接层输入维度计算错误:第二层LSTM输出的是每个样本对应64维隐状态向量,不需要展平操作,直接接输入维度为64的全连接层即可。
- 缺失最后一步将输出乘以400的缩放逻辑。
正确实现代码
import torch import torch.nn as nn import torch.nn.functional as F window_size = 30 batch_size = 32 class TimeSeriesPredictModel(nn.Module): def __init__(self): super().__init__() # 因果卷积:输入通道1(单变量时序),输出通道64,卷积核3 self.conv1d = nn.Conv1d(in_channels=1, out_channels=64, kernel_size=3, stride=1, padding=2) # 第一层LSTM:返回所有时间步输出,对应TF的return_sequences=True self.lstm1 = nn.LSTM(input_size=64, hidden_size=64, batch_first=True) # 第二层LSTM:默认返回所有时间步,后续手动取最后一步输出,对应TF默认return_sequences=False self.lstm2 = nn.LSTM(input_size=64, hidden_size=64, batch_first=True) # 全连接层 self.fc1 = nn.Linear(in_features=64, out_features=30) self.fc2 = nn.Linear(in_features=30, out_features=30) self.fc3 = nn.Linear(in_features=30, out_features=1) def forward(self, x): # x初始形状: (batch_size, window_size, 1) # 转换为Conv1d要求的 (batch, channels, seq_len) 格式 out = x.permute(0, 2, 1) out = F.relu(self.conv1d(out)) # 裁掉卷积右侧多出来的2个时间步,保证序列长度为window_size,实现因果卷积效果 out = out[:, :, :-2] # 转换回LSTM要求的 (batch, seq_len, features) 格式(batch_first=True) out = out.permute(0, 2, 1) out, _ = self.lstm1(out) out, _ = self.lstm2(out) # 取最后一个时间步的输出,形状变为 (batch_size, 64) out = out[:, -1, :] # 全连接前向传播 out = F.relu(self.fc1(out)) out = F.relu(self.fc2(out)) out = self.fc3(out) # 对应TF的Lambda层,输出缩放400倍 out = out * 400 return out # 维度测试 if __name__ == "__main__": model = TimeSeriesPredictModel() # 构造符合要求的测试输入:(32个样本, 窗口长度30, 单特征) test_input = torch.rand((batch_size, window_size, 1), dtype=torch.float32) output = model(test_input) print(f"输入形状: {test_input.shape}, 输出形状: {output.shape}") # 预期打印:输入形状: torch.Size([32, 30, 1]), 输出形状: torch.Size([32, 1])
维度流转对应说明
以batch_size=32为例,每一步张量形状完全对齐原TensorFlow模型:
- 初始输入:
(32, 30, 1) - 卷积处理后:
(32, 30, 64) - 第一层LSTM输出:
(32, 30, 64)(保留所有时间步) - 第二层LSTM取最后一步输出:
(32, 64) - 三层全连接处理后:
(32, 1) - 缩放后最终输出:
(32, 1),即每个样本对应一个下一位置的预测值。
内容的提问来源于stack exchange,提问作者imantha
相关产品推荐
相关产品推荐

