You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Jetson Nano上加速Python cv2.VideoWriter视频录制

解决Jetson Nano上cv2.VideoWriter录制帧率过低的问题

问题根源

你的代码存在两个核心性能瓶颈:

  1. CPU编码:使用cv2.VideoWriter的mp4v编码器是纯CPU运算,Jetson Nano的ARM CPU性能有限,无法高效处理2560x1440分辨率的实时编码。
  2. 不必要的CPU/GPU数据拷贝:捕获pipeline中通过videoconvert将GPU解码后的帧转换为CPU内存中的BGR格式,后续写入又要将CPU帧传回编码,两次数据拷贝+CPU运算严重拖慢帧率。

优化方案:基于Gstreamer硬件加速的端到端录制

利用Jetson平台的NVIDIA硬件编码组件nvv4l2h264enc,结合Gstreamer pipeline实现全GPU加速的视频捕获与编码,避免CPU参与密集运算。

优化后的代码

import os
import time
import cv2

width = 2560
height = 1440
framerate = 30
video_path = 'fps_test.mp4'

# 捕获pipeline:保留GPU解码后的BGRx格式,避免CPU端颜色转换
capture_pipeline = f"v4l2src device=/dev/video0 io-mode=2 " \
                   f"! image/jpeg, width={width}, height={height}, framerate={framerate}/1, format=MJPG " \
                   f"! nvv4l2decoder mjpeg=1 " \
                   f"! nvvidconv flip-method=4 " \
                   f"! video/x-raw, format=BGRx " \
                   f"! appsink drop=1"

# 写入pipeline:使用NVIDIA硬件编码器,全GPU流程编码写入
write_pipeline = f"appsrc ! video/x-raw, format=BGRx, width={width}, height={height}, framerate={framerate}/1 " \
                 f"! nvvidconv " \
                 f"! nvv4l2h264enc bitrate=8000000 preset-level=2 " \
                 f"! h264parse " \
                 f"! mp4mux " \
                 f"! filesink location={video_path} sync=0"

print(f"捕获pipeline: gst-launch-1.0 {capture_pipeline}\n")
print(f"写入pipeline: gst-launch-1.0 {write_pipeline}\n")

v_cap = v_writer = None


def main():
    global v_cap, v_writer
    # 初始化Gstreamer后端的VideoWriter,fourcc设为0由pipeline指定编码器
    v_writer = cv2.VideoWriter(write_pipeline, cv2.CAP_GSTREAMER, 0, framerate, (width, height))
    v_cap = cv2.VideoCapture(capture_pipeline, cv2.CAP_GSTREAMER)

    if not v_cap.isOpened() or not v_writer.isOpened():
        print("无法打开捕获设备或写入器")
        return

    prev_frame_time = time.time()
    frame_count = 0
    while v_cap.isOpened():
        ret_val, frame = v_cap.read()
        if not ret_val:
            break

        # 直接写入BGRx格式帧,无需转换
        v_writer.write(frame)

        # 计算平均FPS(减少打印频率,避免影响性能)
        frame_count += 1
        curr_time = time.time()
        if curr_time - prev_frame_time >= 1.0:
            avg_fps = frame_count / (curr_time - prev_frame_time)
            print(f"平均FPS = {avg_fps:.3f}")
            prev_frame_time = curr_time
            frame_count = 0

    v_cap.release()
    v_writer.release()
    cv2.destroyAllWindows()


if __name__ == '__main__':
    try:
        main()
    except KeyboardInterrupt:
        pass
    finally:
        if v_cap is not None:
            v_cap.release()
        if v_writer is not None:
            v_writer.release()
        cv2.destroyAllWindows()
        # 注释掉删除文件的代码,方便验证录制结果
        # if os.path.exists(video_path):
        #     os.remove(video_path)

关键优化点说明

  1. 捕获流程:

    • 移除videoconvert和format=BGR转换,保留GPU解码后的BGRx格式,避免CPU端的颜色转换运算与数据拷贝。
    • nvv4l2decoder:硬件加速MJPEG解码,利用Jetson的GPU资源。
  2. 写入流程:

    • 使用cv2.CAP_GSTREAMER后端初始化VideoWriter,通过自定义pipeline指定硬件编码流程。
    • nvv4l2h264enc:NVIDIA硬件H.264编码器,性能远高于CPU编码,支持2560x1440@30fps的实时编码。
    • sync=0:关闭filesink的同步机制,避免写入阻塞拖慢帧率。
    • bitrate=8000000:设置编码比特率(可根据需求调整,值越大画质越好),preset-level=2平衡编码速度与画质。
  3. 性能优化细节:

    • 改为计算平均FPS并每秒打印一次,避免高频print操作占用CPU资源。
    • 增加设备打开状态检查,提前排查初始化失败问题。

额外优化建议

  • 如果需要对帧进行图像处理,优先使用cv2.cuda系列函数(需确保OpenCV编译时开启CUDA支持),在GPU上完成处理,避免数据拷贝到CPU。
  • 若不需要Python介入处理帧,可直接使用Gstreamer命令行完成录制,性能更优:
gst-launch-1.0 v4l2src device=/dev/video0 io-mode=2 ! image/jpeg, width=2560, height=1440, framerate=30/1, format=MJPG ! nvv4l2decoder mjpeg=1 ! nvvidconv flip-method=4 ! nvv4l2h264enc bitrate=8000000 preset-level=2 ! h264parse ! mp4mux ! filesink location=fps_test.mp4 sync=0

内容的提问来源于stack exchange,提问作者Jet C.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 05:51:18