使用RTSP流时TensorFlow目标检测卡顿延迟问题求助
问题:RTSP流TensorFlow目标检测延迟卡顿(无GPU PC)
我按照示例实现了USB摄像头的TensorFlow目标检测功能,但切换为IP摄像头的H264编码RTSP流后,出现约30秒延迟且画面时常卡顿。以下是核心处理代码:
import cv2 cap = cv2.VideoCapture("rtsp://192.168.200.1:5544/stream1") # Running the tensorflow session with detection_graph.as_default(): with tf.Session(graph=detection_graph) as sess: ret = True while (ret): ret,image_np = cap.read() # Expand dimensions since the model expects images to have shape: [1, None, None, 3] image_np_expanded = np.expand_dims(image_np, axis=0) image_tensor = detection_graph.get_tensor_by_name('image_tensor:0') # Each box represents a part of the image where a particular object was detected. boxes = detection_graph.get_tensor_by_name('detection_boxes:0') # Each score represent how level of confidence for each of the objects. # Score is shown on the result image, together with the class label. scores = detection_graph.get_tensor_by_name('detection_scores:0') classes = detection_graph.get_tensor_by_name('detection_classes:0') num_detections = detection_graph.get_tensor_by_name('num_detections:0') # Actual detection. (boxes, scores, classes, num_detections) = sess.run( [boxes, scores, classes, num_detections], feed_dict={image_tensor: image_np_expanded}) # Visualization of the results of a detection. vis_util.visualize_boxes_and_labels_on_image_array( image_np, np.squeeze(boxes), np.squeeze(classes).astype(np.int32), np.squeeze(scores), category_index, use_normalized_coordinates=True, line_thickness=8) # plt.figure(figsize=IMAGE_SIZE) # plt.imshow(image_np) cv2.imshow('image',cv2.resize(image_np,(1280,960))) if cv2.waitKey(25) & 0xFF == ord('q'): cv2.destroyAllWindows() cap.release() break
我是Python和TensorFlow新手,在无GPU的PC上,是否需要修改代码以适配RTSP流?
回答
当然需要调整代码来适配RTSP流!RTSP和USB摄像头的帧传输、缓存逻辑差异很大,加上你没有GPU,推理速度本身受限,这两个因素叠加就导致了延迟和卡顿。下面是几个针对性的优化点,都是小改动就能见效:
1. 优化OpenCV的RTSP读取参数
OpenCV默认会给RTSP流设置较大的缓存,用来应对网络波动,但这会导致严重延迟。你可以在初始化VideoCapture后加上这两行:
cap = cv2.VideoCapture("rtsp://192.168.200.1:5544/stream1") # 把缓存大小设为1,只保留最新一帧 cap.set(cv2.CAP_PROP_BUFFERSIZE, 1) # 强制使用TCP传输(UDP容易丢包,导致画面卡顿) cap.set(cv2.CAP_PROP_RTSP_TRANSPORT, cv2.CAP_RTSP_TRANSPORT_TCP)
这一步能直接砍掉大部分由缓存带来的延迟。
2. 跳过旧帧,只处理最新帧
你的检测推理速度(无GPU下)大概率跟不上RTSP流的帧率,导致旧帧堆积,延迟越来越高。可以在循环里先快速抓取所有缓存的帧,只解码最新的那一帧:
while (ret): # 先快速抓取所有缓存的旧帧,只留最后一帧处理 for _ in range(int(cap.get(cv2.CAP_PROP_FPS)) // 2): cap.grab() # grab()只抓帧不解码,比read()快很多 ret, image_np = cap.retrieve() # 解码最新的帧 # 后面的检测逻辑保持不变...
这样你就不会再处理几秒钟前的旧画面,延迟会显著降低。
3. 优化TensorFlow推理速度(无GPU重点)
没有GPU的话,CPU推理速度是瓶颈,得从模型和输入尺寸入手:
- 换轻量模型:比如把当前模型换成SSD MobileNet v2,它的推理速度比Faster R-CNN快好几倍,精度也足够日常使用。
- 缩小输入图像尺寸:在把图像输入模型前,先缩小尺寸,减少计算量:
# 在expand_dims之前先resize image_np = cv2.resize(image_np, (640, 480)) # 可以根据需求调整,比如480x360更快 image_np_expanded = np.expand_dims(image_np, axis=0) - 简化可视化:把
line_thickness改小(比如设为2),或者暂时关闭可视化,先测试推理速度,确认是可视化拖慢还是推理拖慢。
4. 调整waitKey的时间
你的代码里用了cv2.waitKey(25),对应40fps的刷新频率,但如果你的推理速度达不到这么快,这个时间可以适当调小,比如改成10ms,不过前面的优化才是重点。
按照这些步骤改完,延迟应该能降到1-2秒以内,卡顿也会明显减少。
内容的提问来源于stack exchange,提问作者Harry Boy
相关产品推荐
相关产品推荐

