You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ffmpeg-python输出与图像尺寸不符,能否直接生成numpy数组?

Directly Get Numpy Array Frames with ffmpeg-python (No OpenCV Needed)

Absolutely—you can tweak your ffmpeg-python pipeline to output raw, uncompressed pixel data directly, which we can convert to a numpy array without relying on OpenCV. Here's how to do it, tailored to your 1920x1080 YUV420p video:

Option 1: Output RGB24 for Easy Numpy Handling

RGB24 is a packed pixel format (each pixel uses 3 consecutive bytes for R, G, B) that maps cleanly to numpy's 3-channel array structure. This is great if you need RGB data for further processing:

import ffmpeg
import numpy as np

in_filename = "your_input_video.mp4"
frame_num = 0  # Replace with your target frame number

# Fetch video metadata to avoid hardcoding dimensions (optional but robust)
probe = ffmpeg.probe(in_filename)
video_stream = next((stream for stream in probe['streams'] if stream['codec_type'] == 'video'), None)
width = int(video_stream['width'])
height = int(video_stream['height'])

# Extract raw RGB24 frame via ffmpeg
raw_frame_bytes, _ = (
    ffmpeg
    .input(in_filename)
    .filter('select', f'gte(n,{frame_num})')  # Select the target frame
    .output(
        'pipe:',
        vframes=1,          # Only output 1 frame
        format='rawvideo',  # Output uncompressed raw pixels
        pix_fmt='rgb24'     # Use RGB24 pixel format
    )
    .run(capture_stdout=True, quiet=True)  # Capture stdout, suppress logs
)

# Convert raw bytes to numpy array
rgb_frame = np.frombuffer(raw_frame_bytes, dtype=np.uint8).reshape((height, width, 3))

print(rgb_frame.shape)  # Should output (1080, 1920, 3)

Option 2: Output YUV420p Directly (For Performance)

If you want to skip color conversion (to save CPU on embedded systems), you can output the original YUV420p format. Note that YUV420p is planar, so it's split into three separate planes:

  • Y plane: Full resolution (1920x1080)
  • U/V planes: Half resolution (960x540 each)

Here's how to parse it into numpy arrays:

import ffmpeg
import numpy as np

in_filename = "your_input_video.mp4"
frame_num = 0

# Get video dimensions
probe = ffmpeg.probe(in_filename)
video_stream = next((stream for stream in probe['streams'] if stream['codec_type'] == 'video'), None)
width = int(video_stream['width'])
height = int(video_stream['height'])

# Extract raw YUV420p frame
raw_yuv_bytes, _ = (
    ffmpeg
    .input(in_filename)
    .filter('select', f'gte(n,{frame_num})')
    .output(
        'pipe:',
        vframes=1,
        format='rawvideo',
        pix_fmt='yuv420p'  # Match your video's pixel format
    )
    .run(capture_stdout=True, quiet=True)
)

# Split into Y, U, V planes
y_size = width * height
uv_size = (width // 2) * (height // 2)

y_plane = np.frombuffer(raw_yuv_bytes[:y_size], dtype=np.uint8).reshape((height, width))
u_plane = np.frombuffer(raw_yuv_bytes[y_size:y_size+uv_size], dtype=np.uint8).reshape((height//2, width//2))
v_plane = np.frombuffer(raw_yuv_bytes[y_size+uv_size:], dtype=np.uint8).reshape((height//2, width//2))

# Use the planes directly, or implement a lightweight YUV-to-RGB conversion if needed

Key Notes:

  • format='rawvideo': Tells ffmpeg to output uncompressed pixel data instead of encoded formats like JPEG. This eliminates the need for decoding libraries like OpenCV.
  • quiet=True: Suppresses ffmpeg's default log output, which is helpful for embedded systems where you don't want extra console noise.
  • Metadata Probing: Using ffmpeg.probe() to get video dimensions ensures your code works with any video size, not just 1920x1080.

内容的提问来源于stack exchange,提问作者Lionnel104

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:33:05