You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用torchvision保存灰度视频?维度不匹配问题求助

问题:torchvision保存灰度视频时维度不匹配报错

问题详情

使用torchvision的write_video方法保存灰度视频时出现维度不匹配错误,处理彩色视频时可正常运行。报错核心信息:ValueError: Expected numpy array with ndim '3' but got '2'

完整报错堆栈

The detector is mediapipe
INFO: Created TensorFlow Lite XNNPACK delegate for CPU.
/Users/shaimaa/Downloads/LIP_Reading/Code/audio-visual-dataset-master/results_news/news/LRWAR_new/Politics_news_18_p2_170.mp4
this is data_filename : Politics_news_18_p2_170.mp4
this is dst_filename : /Users/shaimaa/Downloads/LIP_Reading/Code/audio-visual-dataset-master/results_news/news/LRWAR_Mouth/Politics_news_18_p2_170.mp4
/Users/shaimaa/Downloads/LIP_Reading/Code/audio-visual-dataset-master/results_news/news/LRWAR_new/Politics_news_18_p2_170.mp4
[None, None, None, None, None, None, None, None, None, None]
[array([[ 75,  72],
       [151,  71],
       [ 99, 112],
       [105, 153]]), array([[ 74,  73],
       [151,  70],
       [ 99, 112],
       [106, 153]]), array([[ 76,  71],
       [154,  70],
       [103, 114],
       [109, 154]]), array([[ 77,  70],
       [155,  69],
       [103, 112],
       [109, 152]]), array([[ 76,  70],
       [154,  68],
       [102, 112],
       [108, 152]]), array([[ 76,  68],
       [152,  67],
       [100, 108],
       [106, 148]]), array([[ 75,  69],
       [149,  68],
       [ 99, 107],
       [105, 147]]), array([[ 74,  70],
       [149,  71],
       [ 97, 109],
       [102, 149]]), array([[ 75,  72],
       [151,  71],
       [ 98, 110],
       [104, 152]]), array([[ 72,  73],
       [148,  72],
       [ 96, 111],
       [102, 152]])]
The landmarks was saved to /Users/shaimaa/Downloads/LIP_Reading/Code/audio-visual-dataset-master/results_news/news/LRWAR_Mouth/Politics_news_18_p2_170.mp4
torch.Size([10, 96, 96])
Error executing job with overrides: ['data_dir=/Users/shaimaa/Downloads/LIP_Reading/Code/audio-visual-dataset-master/results_news/news/LRWAR_new', 'dst_dir=/Users/shaimaa/Downloads/LIP_Reading/Code/audio-visual-dataset-master/results_news/news/LRWAR_Mouth']
Traceback (most recent call last):
  File "crop_mouth.py", line 63, in <module>
    main()
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/main.py", line 94, in decorated_main
    _run_hydra(
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/_internal/utils.py", line 394, in _run_hydra
    _run_app(
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/_internal/utils.py", line 457, in _run_app
    run_and_report(
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/_internal/utils.py", line 223, in run_and_report
    raise ex
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/_internal/utils.py", line 220, in run_and_report
    return func()
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/_internal/utils.py", line 458, in <lambda>
    lambda: hydra.run(
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/_internal/hydra.py", line 132, in run
    _ = ret.return_value
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/core/utils.py", line 260, in return_value
    raise self._return_value
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/hydra/core/utils.py", line 186, in run_job
    ret.return_value = task_function(task_cfg)
  File "crop_mouth.py", line 58, in main
    save2vid(dst_filename, data, fps)
  File "crop_mouth.py", line 21, in save2vid
    torchvision.io.write_video(filename, vid, frames_per_second)
  File "/Users/shaimaa/opt/anaconda3/lib/python3.8/site-packages/torchvision/io/video.py", line 132, in write_video
    frame = av.VideoFrame.from_ndarray(img, format="rgb24")
  File "av/video/frame.pyx", line 408, in av.video.frame.VideoFrame.from_ndarray
  File "av/utils.pyx", line 72, in av.utils.check_ndarray
ValueError: Expected numpy array with ndim `3` but got `2`

原因分析

从tensor.Size([10, 96, 96])可以看出,灰度视频的张量维度为**[帧数, 高度, 宽度]**,而torchvision.io.write_video要求输入张量为3D(单帧:[高度, 宽度, 通道数])或4D(整段视频:[帧数, 高度, 宽度, 通道数])。灰度视频缺少通道维度(通道数应为1),导致维度不匹配。

解决方案

修改save2vid函数,为灰度视频补充通道维度,确保输入符合write_video的要求,以下两种方式可选:

方式1:转为RGB格式保存(兼容性强)

将灰度单通道复制为3通道,模拟RGB格式,适配大多数视频编码器:

def save2vid(filename, vid, fps):
    # 检查是否缺少通道维度
    if vid.dim() == 3:
        # 添加通道维度并复制为3通道
        vid = vid.unsqueeze(-1).repeat(1, 1, 1, 3)
    torchvision.io.write_video(filename, vid, fps)

方式2:保存原生灰度视频(推荐,文件更小)

添加单通道维度,并指定灰度格式写入:

def save2vid(filename, vid, fps):
    # 检查是否缺少通道维度
    if vid.dim() == 3:
        # 添加单通道维度
        vid = vid.unsqueeze(-1)
    # 指定灰度格式写入视频
    torchvision.io.write_video(filename, vid, fps, format="gray")

注意:部分视频编码器不支持灰度格式,若出现编码错误,优先使用方式1。


内容的提问来源于stack exchange,提问作者Shema

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 02:47:07