You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

视频暴力检测模型RuntimeError:输入通道数不匹配问题求助

问题分析与解决方案

错误根源

报错RuntimeError: Given groups=1, weight of size [64, 3, 3, 7, 7], expected input[1, 8, 3, 112, 112] to have 3 channels, but got 8 channels instead的核心原因是输入张量维度顺序与R3D-18的要求不匹配,且错误地将无时间维度的单帧输入给了需要时间维度的3D卷积模型:

  • R3D-18作为视频3D卷积模型,要求输入格式为(batch_size, channels, num_frames, height, width)
  • 你的代码在forward函数中错误地对输入做了维度置换,后续循环输入单帧的操作,导致模型将视频帧的数量误判为通道数,触发通道不匹配错误。

解决方案

方案1:正确使用R3D提取视频序列特征(推荐)

R3D-18本身就是为处理视频序列设计的,无需循环处理单帧,直接输入整个视频片段即可提取序列特征,再传入LSTM。修改模型的forward函数如下:

def forward(self, x):
    print("\n--- Forward Started ---")
    print("Input Shape:", x.shape)  # (batch, 3, frames, 112, 112),符合R3D输入要求

    # 直接用R3D处理整个视频片段,输出形状为(batch, 512)
    cnn_out = self.cnn(x)
    # 调整形状为(batch, 1, 512),适配LSTM的batch_first=True格式
    cnn_features = cnn_out.unsqueeze(1)
    print("LSTM Input:", cnn_features.shape)  # (batch, 1, 512)

    lstm_out, _ = self.lstm(cnn_features)
    lstm_out = lstm_out[:, -1, :] 
    output = self.fc(lstm_out)

    print("Model Output:", output.shape)  # (batch, 1)
    print("--- Forward Finished ---\n")

    return output

方案2:若需提取单帧特征(调整输入维度)

如果确实需要提取每帧特征,建议改用2D卷积模型(如ResNet18);若坚持使用R3D,需给单帧添加时间维度(将单帧视为长度为1的视频序列),修改forward函数的循环部分:

def forward(self, x):
    print("\n--- Forward Started ---")
    print("Input Shape:", x.shape)  # (batch, 3, frames, 112, 112)

    # 调整为(batch, frames, 3, 112, 112),方便循环取单帧
    x = x.permute(0, 2, 1, 3, 4)
    print("Permute:", x.shape)  # (batch, frames, 3, 112, 112)

    cnn_features = []
    for t in range(x.shape[1]):  
        frame = x[:, t, :, :, :]  # (batch, 3, 112, 112)
        # 给单帧添加时间维度,变成(batch, 3, 1, 112, 112),符合R3D输入格式
        frame = frame.unsqueeze(2)
        cnn_out = self.cnn(frame) # (batch, 512)
        cnn_features.append(cnn_out.unsqueeze(1))  # (batch, 1, 512) 
    
    cnn_features = torch.cat(cnn_features, dim=1)
    print("LSTM Input:", cnn_features.shape)  # (batch, frames, 512)

    lstm_out, _ = self.lstm(cnn_features)
    lstm_out = lstm_out[:, -1, :] 
    output = self.fc(lstm_out)

    print("Model Output:", output.shape)  # (batch, 1)
    print("--- Forward Finished ---\n")

    return output

额外说明

  • 你的数据集输出和DataLoader的维度是正确的:(batch, 3, frames, 112, 112),完全符合R3D的输入要求,方案1中无需额外维度置换操作。
  • 方案1的运行效率远高于方案2,因为R3D可以并行处理整个视频序列,避免了循环的开销。

内容的提问来源于stack exchange,提问作者Can Gürcüoğlu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 02:54:54