PyTorch自动编码器图像压缩报错:mat1与mat2形状无法相乘
问题分析与解决
报错原因
报错mat1 and mat2 shapes cannot be multiplied (720x120 and 43200x512)的核心是输入张量与全连接层的形状不匹配:
- 你定义的第一层全连接层
Linear(3*120*120, 512),要求每个输入样本是43200维向量(对应3×120×120图像展平后的尺寸)。 - 但实际输入到该层的张量形状是
(720, 120),即每个样本仅120维,完全不符合层的输入要求。
你提到做了展平操作,大概率是展平方式错误——要么没保留批次维度,要么展平的维度顺序不对,甚至输入图像的维度格式本身就不符合PyTorch的标准要求。
解决方案
1. 正确在模型内完成展平
修改forward函数,从第1维度开始展平(第0维度是批次维度,必须保留),确保每个样本被展平为43200维。同时,解码后可以将向量还原为原始图像形状:
def forward(self, x): # 展平通道在前的图像张量:(batch_size,3,120,120) → (batch_size,43200) x = x.flatten(start_dim=1) encoded = self.encoder(x) decoded = self.decoder(encoded) # 将解码后的向量还原为图像形状 decoded = decoded.reshape(-1, 3, 120, 120) return decoded
2. 确认输入图像的维度格式
确保输入张量是PyTorch标准的通道在前格式:(batch_size, 3, 120, 120)。如果你的数据是通道在后(比如(batch_size, 120, 120, 3)),需要先转置维度:
# 输入为通道在后格式时,先转成通道在前 x = x.permute(0, 3, 1, 2)
完整修改后的代码
class AE(torch.nn.Module): def __init__(self): super().__init__() self.encoder = torch.nn.Sequential( torch.nn.Linear(3*120*120, 512), torch.nn.ReLU(), torch.nn.Linear(512, 256), torch.nn.ReLU(), torch.nn.Linear(256, 128), torch.nn.ReLU(), torch.nn.Linear(128, 120) ) self.decoder = torch.nn.Sequential( torch.nn.Linear(120, 128), torch.nn.ReLU(), torch.nn.Linear(128, 256), torch.nn.ReLU(), torch.nn.Linear(256, 512), torch.nn.ReLU(), torch.nn.Linear(512, 3*120*120), torch.nn.Sigmoid() ) def forward(self, x): # 通道在前格式下展平 x = x.flatten(start_dim=1) encoded = self.encoder(x) decoded = self.decoder(encoded) # 将解码后的向量还原为图像形状 decoded = decoded.reshape(-1, 3, 120, 120) return decoded
额外说明
- 如果展平操作是在模型外完成的,要确保最终输入张量形状是
(batch_size, 43200),而非其他错误形状。 - 报错中的
720是你的批次大小,120是错误的样本特征数,说明展平时误把图像的某一维度当成了样本维度,导致形状完全混乱。
内容的提问来源于stack exchange,提问作者Andrea
相关产品推荐
相关产品推荐

