OpenCV搭配RealSense运行目标检测时帧率骤降问题求助
RealSense D455集成目标检测距离计算功能帧率异常下降问题
问题现象
- 硬件采用RealSense D455-C深度相机,在OpenCV目标检测示例基础上新增基于深度帧的距离计算逻辑后出现帧率暴跌
- 单独运行相机彩色流、深度流采集程序可稳定保持30fps,启动目标检测推理后帧率仅为5-8fps
- 检索公开技术资料未找到有效解决方案
复现代码
import pyrealsense2 as rs import numpy as np import cv2 import tensorflow as tf # 配置深度流、彩色流参数 pipeline = rs.pipeline() config = rs.config() config.enable_stream(rs.stream.color, 1280, 720, rs.format.bgr8, 30) dec_filter = rs.decimation_filter() print("[INFO] Starting streaming...") config.enable_stream(rs.stream.depth, 1280, 720, rs.format.z16, 30) pipeline.start(config) print("[INFO] Camera ready.") print("[INFO] Loading model...") PATH_TO_CKPT = "C:/Testing Camera/faster_rcnn_inception_v2_coco_2018_01_28/frozen_inference_graph.pb" # 加载TensorFlow模型到内存 detection_graph = tf.Graph() with detection_graph.as_default(): od_graph_def = tf.compat.v1.GraphDef() with tf.compat.v1.gfile.GFile(PATH_TO_CKPT, 'rb') as fid: serialized_graph = fid.read() od_graph_def.ParseFromString(serialized_graph) tf.compat.v1.import_graph_def(od_graph_def, name='') sess = tf.compat.v1.Session(graph=detection_graph) # 定义输入输出张量 image_tensor = detection_graph.get_tensor_by_name('image_tensor:0') detection_boxes = detection_graph.get_tensor_by_name('detection_boxes:0') detection_scores = detection_graph.get_tensor_by_name('detection_scores:0') detection_classes = detection_graph.get_tensor_by_name('detection_classes:0') num_detections = detection_graph.get_tensor_by_name('num_detections:0') print("[INFO] Model loaded.") colors_hash = {} while True: frames = pipeline.wait_for_frames() color_frame = frames.get_color_frame() depth_frame = frames.get_depth_frame() # 转换帧为numpy数组 color_image = np.asanyarray(color_frame.get_data()) scaled_size = (color_frame.width, color_frame.height) # 扩展维度匹配模型输入要求 [1, None, None, 3] image_expanded = np.expand_dims(color_image, axis=0) # 执行推理 (boxes, scores, classes, num) = sess.run([detection_boxes, detection_scores, detection_classes, num_detections], feed_dict={image_tensor: image_expanded}) boxes = np.squeeze(boxes) classes = np.squeeze(classes).astype(np.int32) scores = np.squeeze(scores) depth_image = np.asanyarray(depth_frame.get_data()) for idx in range(int(num)): class_ = classes[idx] score = scores[idx] box = boxes[idx] if class_ not in colors_hash: colors_hash[class_] = tuple(np.random.choice(range(256), size=3)) if score > 0.6: #检测置信度阈值,默认0.6 left = int(box[4] * color_frame.width)#x top = int(box[0] * color_frame.height)#y right = int(box[4] * color_frame.width) bottom = int(box[4] * color_frame.height) p1 = (left, top) p2 = (right, bottom) # 绘制检测框 r, g, b = colors_hash[class_] cv2.rectangle(color_image, p1, p2, (int(r), int(g), int(b)), 2, 1) y = int(top/2) x = int(left/2) print("X: "+str(left)+" Y: "+str(top)+" Distance: "+str(depth_image[y,x]))#深度帧索引顺序为y在前x在后 cv2.namedWindow('RealSense', cv2.WINDOW_AUTOSIZE) cv2.imshow('RealSense', color_image) cv2.waitKey(1) print("[INFO] stop streaming ...") pipeline.stop()
运行日志
2022-06-14 12:52:14.058976: I tensorflow/stream_executor/cuda/cudart_stub.cc:29] Ignore above cudart dlerror if you do not have a GPU set up on your machine. [INFO] Starting streaming... [INFO] Camera ready. [INFO] Loading model... 2022-06-14 12:52:18.251559: W tensorflow/stream_executor/platform/default/dso_loader.cc:64] Could not load dynamic library 'nvcuda.dll'; dlerror: nvcuda.dll not found 2022-06-14 12:52:18.251871: W tensorflow/stream_executor/cuda/cuda_driver.cc:269] failed call to cuInit: UNKNOWN ERROR (303) 2022-06-14 12:52:18.257821: I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:169] retrieving CUDA diagnostic information for host: E-5CG1168PR1 2022-06-14 12:52:18.258228: I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:176] hostname: E-5CG1168PR1 2022-06-14 12:52:18.258624: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX AVX2 To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags. [INFO] Model loaded. 2022-06-14 12:52:18.431314: I tensorflow/compiler/mlir/mlir_graph_optimization_pass.cc:354] MLIR V1 optimization pass is not enabled 2022-06-14 12:52:20.576233: W tensorflow/core/grappler/costs/op_level_cost_estimator.cc:690] Error in PredictCost() for the op: op: "CropAndResize" attr { key: "T" value { type: DT_FLOAT } } attr { key: "extrapolation_value" value { f: 0 } } attr { key: "method" value { s: "bilinear" } } inputs { dtype: DT_FLOAT shape { dim { size: -7 } dim { size: -10 } dim { size: -12 } dim { size: 576 } } } inputs { dtype: DT_FLOAT shape { dim { size: -33 } dim { size: 4 } } } inputs { dtype: DT_INT32 shape { dim { size: -33 } } } inputs { dtype: DT_INT32 shape { dim { size: 2 } } value { dtype: DT_INT32 tensor_shape { dim { size: 2 } } int_val: 14 } } device { type: "CPU" vendor: "GenuineIntel" model: "110" frequency: 2304 num_cores: 8 environment { key: "cpu_instruction_set" value: "SSE, SSE2" } environment { key: "eigen" value: "3.4.90" } l1_cache_size: 32768 l2_cache_size: 262144 l3_cache_size: 8388608 memory_size: 268435456 } outputs { dtype: DT_FLOAT shape { dim { size: -33 } dim { size: 14 } dim { size: 14 } dim { size: 576 } } }
根因分析
- 日志明确显示
nvcuda.dll加载失败,CUDA初始化未完成,TensorFlow全程跑在CPU上做推理,没有调用GPU加速 - 当前使用的
faster_rcnn_inception_v2_coco是两阶段检测模型,本身计算量极大,直接输入1280*720分辨率的原图做CPU推理,单帧耗时本身就会达到120-200ms,对应帧率就是5-8fps,和观测到的现象完全吻合,帧率下降和深度距离计算逻辑无关 - 代码存在多处冗余开销和逻辑bug,会额外拉低帧率、导致结果错误:
- 每帧循环都重复创建
cv2.namedWindow窗口,无意义消耗资源 - 循环内逐检测结果调用
print输出,频繁IO操作会阻塞主线程 - 边界框坐标索引完全写错,Faster RCNN输出的box顺序为
[ymin, xmin, ymax, xmax],代码里left、right、bottom都错误取了box[4],会导致绘制的检测框完全失效 - 深度值采样点直接对坐标除以2,没有做彩色帧和深度帧的对齐,取到的距离值完全不对应目标位置
- 每帧循环都重复创建
优化方案
- 优先启用GPU推理:安装和TensorFlow版本匹配的CUDA、cuDNN依赖,确认日志里没有CUDA加载失败的报错,让模型推理跑在GPU上,这是帧率提升最核心的手段,1280*720输入下Faster RCNN在普通消费级GPU上就能跑到30fps以上
- 降低推理计算量:
- 不要直接把1280720原始帧喂给模型,先把图像缩放到300300/640*640这类模型常用的推理输入尺寸,推理完成后再把坐标映射回原始分辨率画框,CPU推理速度能提升2-3倍
- 如果不需要极致检测精度,直接替换成SSD-MobileNet、YOLOv5n/YOLOv8n这类轻量单阶段检测模型,CPU上跑720P输入也能轻松达到25fps以上
- 修正代码逻辑问题:
- 把
cv2.namedWindow初始化移到while循环外部,只创建一次窗口 - 去掉循环里的逐帧print输出,需要调试信息可以每隔10帧打印一次,或者把信息绘制到画面上
- 修正边界框坐标取值:
top = int(box[0] * color_frame.height) left = int(box[1] * color_frame.width) bottom = int(box[2] * color_frame.height) right = int(box[3] * color_frame.width) - 新增深度帧对齐逻辑,在Pipeline启动后加对齐配置,把深度帧对齐到彩色帧坐标系再取距离值,避免坐标不匹配:
align = rs.align(rs.stream.color) # 循环内取帧时替换为 frames = pipeline.wait_for_frames() aligned_frames = align.process(frames) color_frame = aligned_frames.get_color_frame() depth_frame = aligned_frames.get_depth_frame() - 距离采样点取检测框中心位置,不要直接用左上角坐标除以2,取值更准确:
depth = depth_frame.get_distance((left+right)//2, (top+bottom)//2),直接调用SDK的get_distance接口会自动返回米为单位的距离值,比自己读depth数组更稳定
- 把
- 可选优化:把相机取流和模型推理拆到两个独立线程,采集线程只负责维护最新的帧缓存,推理线程从缓存取最新帧做检测,避免推理阻塞相机取流导致的帧堆积、延迟升高
内容的提问来源于stack exchange,提问作者D0lan
相关产品推荐
相关产品推荐

