ffmpeg-python输出与图像尺寸不符,能否直接生成numpy数组?
Directly Get Numpy Array Frames with ffmpeg-python (No OpenCV Needed)
Absolutely—you can tweak your ffmpeg-python pipeline to output raw, uncompressed pixel data directly, which we can convert to a numpy array without relying on OpenCV. Here's how to do it, tailored to your 1920x1080 YUV420p video:
Option 1: Output RGB24 for Easy Numpy Handling
RGB24 is a packed pixel format (each pixel uses 3 consecutive bytes for R, G, B) that maps cleanly to numpy's 3-channel array structure. This is great if you need RGB data for further processing:
import ffmpeg import numpy as np in_filename = "your_input_video.mp4" frame_num = 0 # Replace with your target frame number # Fetch video metadata to avoid hardcoding dimensions (optional but robust) probe = ffmpeg.probe(in_filename) video_stream = next((stream for stream in probe['streams'] if stream['codec_type'] == 'video'), None) width = int(video_stream['width']) height = int(video_stream['height']) # Extract raw RGB24 frame via ffmpeg raw_frame_bytes, _ = ( ffmpeg .input(in_filename) .filter('select', f'gte(n,{frame_num})') # Select the target frame .output( 'pipe:', vframes=1, # Only output 1 frame format='rawvideo', # Output uncompressed raw pixels pix_fmt='rgb24' # Use RGB24 pixel format ) .run(capture_stdout=True, quiet=True) # Capture stdout, suppress logs ) # Convert raw bytes to numpy array rgb_frame = np.frombuffer(raw_frame_bytes, dtype=np.uint8).reshape((height, width, 3)) print(rgb_frame.shape) # Should output (1080, 1920, 3)
Option 2: Output YUV420p Directly (For Performance)
If you want to skip color conversion (to save CPU on embedded systems), you can output the original YUV420p format. Note that YUV420p is planar, so it's split into three separate planes:
- Y plane: Full resolution (1920x1080)
- U/V planes: Half resolution (960x540 each)
Here's how to parse it into numpy arrays:
import ffmpeg import numpy as np in_filename = "your_input_video.mp4" frame_num = 0 # Get video dimensions probe = ffmpeg.probe(in_filename) video_stream = next((stream for stream in probe['streams'] if stream['codec_type'] == 'video'), None) width = int(video_stream['width']) height = int(video_stream['height']) # Extract raw YUV420p frame raw_yuv_bytes, _ = ( ffmpeg .input(in_filename) .filter('select', f'gte(n,{frame_num})') .output( 'pipe:', vframes=1, format='rawvideo', pix_fmt='yuv420p' # Match your video's pixel format ) .run(capture_stdout=True, quiet=True) ) # Split into Y, U, V planes y_size = width * height uv_size = (width // 2) * (height // 2) y_plane = np.frombuffer(raw_yuv_bytes[:y_size], dtype=np.uint8).reshape((height, width)) u_plane = np.frombuffer(raw_yuv_bytes[y_size:y_size+uv_size], dtype=np.uint8).reshape((height//2, width//2)) v_plane = np.frombuffer(raw_yuv_bytes[y_size+uv_size:], dtype=np.uint8).reshape((height//2, width//2)) # Use the planes directly, or implement a lightweight YUV-to-RGB conversion if needed
Key Notes:
format='rawvideo': Tells ffmpeg to output uncompressed pixel data instead of encoded formats like JPEG. This eliminates the need for decoding libraries like OpenCV.quiet=True: Suppresses ffmpeg's default log output, which is helpful for embedded systems where you don't want extra console noise.- Metadata Probing: Using
ffmpeg.probe()to get video dimensions ensures your code works with any video size, not just 1920x1080.
内容的提问来源于stack exchange,提问作者Lionnel104
相关产品推荐
相关产品推荐

