如何在Jetson Nano上加速Python cv2.VideoWriter视频录制
解决Jetson Nano上cv2.VideoWriter录制帧率过低的问题
问题根源
你的代码存在两个核心性能瓶颈:
- CPU编码:使用
cv2.VideoWriter的mp4v编码器是纯CPU运算,Jetson Nano的ARM CPU性能有限,无法高效处理2560x1440分辨率的实时编码。 - 不必要的CPU/GPU数据拷贝:捕获pipeline中通过
videoconvert将GPU解码后的帧转换为CPU内存中的BGR格式,后续写入又要将CPU帧传回编码,两次数据拷贝+CPU运算严重拖慢帧率。
优化方案:基于Gstreamer硬件加速的端到端录制
利用Jetson平台的NVIDIA硬件编码组件nvv4l2h264enc,结合Gstreamer pipeline实现全GPU加速的视频捕获与编码,避免CPU参与密集运算。
优化后的代码
import os import time import cv2 width = 2560 height = 1440 framerate = 30 video_path = 'fps_test.mp4' # 捕获pipeline:保留GPU解码后的BGRx格式,避免CPU端颜色转换 capture_pipeline = f"v4l2src device=/dev/video0 io-mode=2 " \ f"! image/jpeg, width={width}, height={height}, framerate={framerate}/1, format=MJPG " \ f"! nvv4l2decoder mjpeg=1 " \ f"! nvvidconv flip-method=4 " \ f"! video/x-raw, format=BGRx " \ f"! appsink drop=1" # 写入pipeline:使用NVIDIA硬件编码器,全GPU流程编码写入 write_pipeline = f"appsrc ! video/x-raw, format=BGRx, width={width}, height={height}, framerate={framerate}/1 " \ f"! nvvidconv " \ f"! nvv4l2h264enc bitrate=8000000 preset-level=2 " \ f"! h264parse " \ f"! mp4mux " \ f"! filesink location={video_path} sync=0" print(f"捕获pipeline: gst-launch-1.0 {capture_pipeline}\n") print(f"写入pipeline: gst-launch-1.0 {write_pipeline}\n") v_cap = v_writer = None def main(): global v_cap, v_writer # 初始化Gstreamer后端的VideoWriter,fourcc设为0由pipeline指定编码器 v_writer = cv2.VideoWriter(write_pipeline, cv2.CAP_GSTREAMER, 0, framerate, (width, height)) v_cap = cv2.VideoCapture(capture_pipeline, cv2.CAP_GSTREAMER) if not v_cap.isOpened() or not v_writer.isOpened(): print("无法打开捕获设备或写入器") return prev_frame_time = time.time() frame_count = 0 while v_cap.isOpened(): ret_val, frame = v_cap.read() if not ret_val: break # 直接写入BGRx格式帧,无需转换 v_writer.write(frame) # 计算平均FPS(减少打印频率,避免影响性能) frame_count += 1 curr_time = time.time() if curr_time - prev_frame_time >= 1.0: avg_fps = frame_count / (curr_time - prev_frame_time) print(f"平均FPS = {avg_fps:.3f}") prev_frame_time = curr_time frame_count = 0 v_cap.release() v_writer.release() cv2.destroyAllWindows() if __name__ == '__main__': try: main() except KeyboardInterrupt: pass finally: if v_cap is not None: v_cap.release() if v_writer is not None: v_writer.release() cv2.destroyAllWindows() # 注释掉删除文件的代码,方便验证录制结果 # if os.path.exists(video_path): # os.remove(video_path)
关键优化点说明
捕获流程:
- 移除
videoconvert和format=BGR转换,保留GPU解码后的BGRx格式,避免CPU端的颜色转换运算与数据拷贝。 nvv4l2decoder:硬件加速MJPEG解码,利用Jetson的GPU资源。
- 移除
写入流程:
- 使用
cv2.CAP_GSTREAMER后端初始化VideoWriter,通过自定义pipeline指定硬件编码流程。 nvv4l2h264enc:NVIDIA硬件H.264编码器,性能远高于CPU编码,支持2560x1440@30fps的实时编码。sync=0:关闭filesink的同步机制,避免写入阻塞拖慢帧率。bitrate=8000000:设置编码比特率(可根据需求调整,值越大画质越好),preset-level=2平衡编码速度与画质。
- 使用
性能优化细节:
- 改为计算平均FPS并每秒打印一次,避免高频
print操作占用CPU资源。 - 增加设备打开状态检查,提前排查初始化失败问题。
- 改为计算平均FPS并每秒打印一次,避免高频
额外优化建议
- 如果需要对帧进行图像处理,优先使用
cv2.cuda系列函数(需确保OpenCV编译时开启CUDA支持),在GPU上完成处理,避免数据拷贝到CPU。 - 若不需要Python介入处理帧,可直接使用Gstreamer命令行完成录制,性能更优:
gst-launch-1.0 v4l2src device=/dev/video0 io-mode=2 ! image/jpeg, width=2560, height=1440, framerate=30/1, format=MJPG ! nvv4l2decoder mjpeg=1 ! nvvidconv flip-method=4 ! nvv4l2h264enc bitrate=8000000 preset-level=2 ! h264parse ! mp4mux ! filesink location=fps_test.mp4 sync=0
内容的提问来源于stack exchange,提问作者Jet C.
相关产品推荐
相关产品推荐

