视频暴力检测模型RuntimeError:输入通道数不匹配问题求助
问题分析与解决方案
错误根源
报错RuntimeError: Given groups=1, weight of size [64, 3, 3, 7, 7], expected input[1, 8, 3, 112, 112] to have 3 channels, but got 8 channels instead的核心原因是输入张量维度顺序与R3D-18的要求不匹配,且错误地将无时间维度的单帧输入给了需要时间维度的3D卷积模型:
- R3D-18作为视频3D卷积模型,要求输入格式为
(batch_size, channels, num_frames, height, width) - 你的代码在
forward函数中错误地对输入做了维度置换,后续循环输入单帧的操作,导致模型将视频帧的数量误判为通道数,触发通道不匹配错误。
解决方案
方案1:正确使用R3D提取视频序列特征(推荐)
R3D-18本身就是为处理视频序列设计的,无需循环处理单帧,直接输入整个视频片段即可提取序列特征,再传入LSTM。修改模型的forward函数如下:
def forward(self, x): print("\n--- Forward Started ---") print("Input Shape:", x.shape) # (batch, 3, frames, 112, 112),符合R3D输入要求 # 直接用R3D处理整个视频片段,输出形状为(batch, 512) cnn_out = self.cnn(x) # 调整形状为(batch, 1, 512),适配LSTM的batch_first=True格式 cnn_features = cnn_out.unsqueeze(1) print("LSTM Input:", cnn_features.shape) # (batch, 1, 512) lstm_out, _ = self.lstm(cnn_features) lstm_out = lstm_out[:, -1, :] output = self.fc(lstm_out) print("Model Output:", output.shape) # (batch, 1) print("--- Forward Finished ---\n") return output
方案2:若需提取单帧特征(调整输入维度)
如果确实需要提取每帧特征,建议改用2D卷积模型(如ResNet18);若坚持使用R3D,需给单帧添加时间维度(将单帧视为长度为1的视频序列),修改forward函数的循环部分:
def forward(self, x): print("\n--- Forward Started ---") print("Input Shape:", x.shape) # (batch, 3, frames, 112, 112) # 调整为(batch, frames, 3, 112, 112),方便循环取单帧 x = x.permute(0, 2, 1, 3, 4) print("Permute:", x.shape) # (batch, frames, 3, 112, 112) cnn_features = [] for t in range(x.shape[1]): frame = x[:, t, :, :, :] # (batch, 3, 112, 112) # 给单帧添加时间维度,变成(batch, 3, 1, 112, 112),符合R3D输入格式 frame = frame.unsqueeze(2) cnn_out = self.cnn(frame) # (batch, 512) cnn_features.append(cnn_out.unsqueeze(1)) # (batch, 1, 512) cnn_features = torch.cat(cnn_features, dim=1) print("LSTM Input:", cnn_features.shape) # (batch, frames, 512) lstm_out, _ = self.lstm(cnn_features) lstm_out = lstm_out[:, -1, :] output = self.fc(lstm_out) print("Model Output:", output.shape) # (batch, 1) print("--- Forward Finished ---\n") return output
额外说明
- 你的数据集输出和DataLoader的维度是正确的:
(batch, 3, frames, 112, 112),完全符合R3D的输入要求,方案1中无需额外维度置换操作。 - 方案1的运行效率远高于方案2,因为R3D可以并行处理整个视频序列,避免了循环的开销。
内容的提问来源于stack exchange,提问作者Can Gürcüoğlu
相关产品推荐
相关产品推荐

