OAK-D部署YOLOv8分割:大列表转numpy数组延迟优化问询
OAK-D部署YOLOv8分割模型:扁平列表转numpy数组性能优化问题
问题描述
我正在用Luxonis的OAK-D设备部署YOLOv8分割模型,当前单帧处理耗时约0.14秒,其中numpy数组转换是核心性能瓶颈:
- 将YOLOv8-seg的output0(长度974400的扁平列表)转换为
(1, 116, 8400)维度numpy数组,耗时0.05秒 - 将output1(长度819200的扁平列表)转换为
(1, 32, 160, 160)维度numpy数组,耗时0.03秒
两者合计占总耗时的57%。我试过np.reshape、np.fromiter、先转np.array再处理的方法,效率都差不多,想找更高效的转换方式。
更新:相关代码
import cv2 import numpy as np import depthai as dai import time from YOLOSeg import YOLOSeg pathYoloBlob = "./yolov8n-seg.blob" # 创建OAK-D管线 pipeline = dai.Pipeline() cam_rgb = pipeline.createColorCamera() cam_rgb.setPreviewSize(640, 640) cam_rgb.setInterleaved(False) nn = pipeline.create(dai.node.NeuralNetwork) nn.setBlobPath(pathYoloBlob) cam_rgb.preview.link(nn.input) xout_rgb = pipeline.createXLinkOut() xout_rgb.setStreamName("rgb") cam_rgb.preview.link(xout_rgb.input) xout_nn_yolo = pipeline.createXLinkOut() xout_nn_yolo.setStreamName("nn_yolo") nn.out.link(xout_nn_yolo.input) # 启动应用 with depthai.Device(pipeline) as device: q_rgb = device.getOutputQueue("rgb") q_nn_yolo = device.getOutputQueue("nn_yolo") frame = None # 归一化检测框坐标到图像尺寸 def frameNorm(frame, bbox): normVals = np.full(len(bbox), frame.shape[0]) normVals[::2] = frame.shape[1] return (np.clip(np.array(bbox), 0, 1) * normVals).astype(int) # 主机端主循环 while True: in_rgb = q_rgb.tryGet() in_nn_yolo = q_nn_yolo.tryGet() if in_rgb is not None: frame = in_rgb.getCvFrame() if in_nn_yolo is not None: # 性能瓶颈所在 output0 = np.reshape(in_nn_yolo.getLayerFp16("output0"), newshape=([1, 116, 8400])) output1 = np.reshape(in_nn_yolo.getLayerFp16("output1"), newshape=([1, 32, 160, 160])) # 获取到两个输出后计算掩码 if len(output0) > 0 and len(output1) > 0: # 后处理部分速度正常 yoloseg = YOLOSeg("", conf_thres=0.3, iou_thres=0.5) yoloseg.prepare_input_for_oakd(frame.shape[:2]) yoloseg.segment_objects_from_oakd(output0,output1) combined_img = yoloseg.draw_masks(frame.copy()) cv2.imshow("Output", combined_img) else: print("in_nn_yolo EMPTY") else: print("in_rgb EMPTY") # 按q退出循环 if cv2.waitKey(1) == ord('q'): break
优化方案
1. 用底层缓冲区直接转numpy数组(核心优化)
getLayerFp16()会把设备返回的数据转换成Python列表,这一步是性能损耗的关键。改用getLayerBuffer()获取原始内存缓冲区,直接通过np.frombuffer()创建numpy数组,避免中间的Python对象转换:
# 替换原有的reshape代码 output0_buf = in_nn_yolo.getLayerBuffer("output0") output0 = np.frombuffer(output0_buf, dtype=np.float16).reshape((1, 116, 8400)) output1_buf = in_nn_yolo.getLayerBuffer("output1") output1 = np.frombuffer(output1_buf, dtype=np.float16).reshape((1, 32, 160, 160))
这种方式相当于直接映射设备内存到numpy数组,无需逐个元素拷贝,能大幅降低转换耗时。
2. 提前初始化数组复用内存
循环内反复创建numpy数组会带来内存分配开销,可以在循环外提前创建固定形状的空数组,每次处理时直接拷贝数据:
# 循环外初始化固定形状的数组 output0 = np.empty((1, 116, 8400), dtype=np.float16) output1 = np.empty((1, 32, 160, 160), dtype=np.float16) # 循环内替换为: output0_buf = in_nn_yolo.getLayerBuffer("output0") np.copyto(output0, np.frombuffer(output0_buf, dtype=np.float16).reshape(output0.shape)) output1_buf = in_nn_yolo.getLayerBuffer("output1") np.copyto(output1, np.frombuffer(output1_buf, dtype=np.float16).reshape(output1.shape))
3. 转换模型时指定输出维度
如果使用OpenVINO转换blob文件,可通过--output_shape参数直接指定输出为目标维度,让设备端输出结构化张量,主机端无需额外reshape。例如转换命令:
mo.py --input_model yolov8n-seg.onnx --output_shape "[1,116,8400],[1,32,160,160]"
4. 减少不必要的对象创建
代码中每次循环都创建YOLOSeg实例,可移到循环外复用:
# 循环外初始化 yoloseg = YOLOSeg("", conf_thres=0.3, iou_thres=0.5) # 循环内仅调用方法 yoloseg.prepare_input_for_oakd(frame.shape[:2]) yoloseg.segment_objects_from_oakd(output0,output1)
内容的提问来源于stack exchange,提问作者Pedro
相关产品推荐
相关产品推荐

