You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RTSP流时TensorFlow目标检测卡顿延迟问题求助

问题:RTSP流TensorFlow目标检测延迟卡顿(无GPU PC)

我按照示例实现了USB摄像头的TensorFlow目标检测功能,但切换为IP摄像头的H264编码RTSP流后,出现约30秒延迟且画面时常卡顿。以下是核心处理代码:

import cv2
cap = cv2.VideoCapture("rtsp://192.168.200.1:5544/stream1")
# Running the tensorflow session
with detection_graph.as_default():
    with tf.Session(graph=detection_graph) as sess:
        ret = True
        while (ret):
            ret,image_np = cap.read()
            # Expand dimensions since the model expects images to have shape: [1, None, None, 3]
            image_np_expanded = np.expand_dims(image_np, axis=0)
            image_tensor = detection_graph.get_tensor_by_name('image_tensor:0')
            # Each box represents a part of the image where a particular object was detected.
            boxes = detection_graph.get_tensor_by_name('detection_boxes:0')
            # Each score represent how level of confidence for each of the objects.
            # Score is shown on the result image, together with the class label.
            scores = detection_graph.get_tensor_by_name('detection_scores:0')
            classes = detection_graph.get_tensor_by_name('detection_classes:0')
            num_detections = detection_graph.get_tensor_by_name('num_detections:0')
            # Actual detection.
            (boxes, scores, classes, num_detections) = sess.run(
                [boxes, scores, classes, num_detections],
                feed_dict={image_tensor: image_np_expanded})
            # Visualization of the results of a detection.
            vis_util.visualize_boxes_and_labels_on_image_array(
                image_np,
                np.squeeze(boxes),
                np.squeeze(classes).astype(np.int32),
                np.squeeze(scores),
                category_index,
                use_normalized_coordinates=True,
                line_thickness=8)
            # plt.figure(figsize=IMAGE_SIZE)
            # plt.imshow(image_np)
            cv2.imshow('image',cv2.resize(image_np,(1280,960)))
            if cv2.waitKey(25) & 0xFF == ord('q'):
                cv2.destroyAllWindows()
                cap.release()
                break

我是Python和TensorFlow新手,在无GPU的PC上,是否需要修改代码以适配RTSP流?


回答

当然需要调整代码来适配RTSP流!RTSP和USB摄像头的帧传输、缓存逻辑差异很大,加上你没有GPU,推理速度本身受限,这两个因素叠加就导致了延迟和卡顿。下面是几个针对性的优化点,都是小改动就能见效:

1. 优化OpenCV的RTSP读取参数

OpenCV默认会给RTSP流设置较大的缓存,用来应对网络波动,但这会导致严重延迟。你可以在初始化VideoCapture后加上这两行:

cap = cv2.VideoCapture("rtsp://192.168.200.1:5544/stream1")
# 把缓存大小设为1,只保留最新一帧
cap.set(cv2.CAP_PROP_BUFFERSIZE, 1)
# 强制使用TCP传输(UDP容易丢包,导致画面卡顿)
cap.set(cv2.CAP_PROP_RTSP_TRANSPORT, cv2.CAP_RTSP_TRANSPORT_TCP)

这一步能直接砍掉大部分由缓存带来的延迟。

2. 跳过旧帧,只处理最新帧

你的检测推理速度(无GPU下)大概率跟不上RTSP流的帧率,导致旧帧堆积,延迟越来越高。可以在循环里先快速抓取所有缓存的帧,只解码最新的那一帧:

while (ret):
    # 先快速抓取所有缓存的旧帧,只留最后一帧处理
    for _ in range(int(cap.get(cv2.CAP_PROP_FPS)) // 2):
        cap.grab()  # grab()只抓帧不解码,比read()快很多
    ret, image_np = cap.retrieve()  # 解码最新的帧
    # 后面的检测逻辑保持不变...

这样你就不会再处理几秒钟前的旧画面,延迟会显著降低。

3. 优化TensorFlow推理速度(无GPU重点)

没有GPU的话,CPU推理速度是瓶颈,得从模型和输入尺寸入手:

  • 换轻量模型:比如把当前模型换成SSD MobileNet v2,它的推理速度比Faster R-CNN快好几倍,精度也足够日常使用。
  • 缩小输入图像尺寸:在把图像输入模型前,先缩小尺寸,减少计算量:
    # 在expand_dims之前先resize
    image_np = cv2.resize(image_np, (640, 480))  # 可以根据需求调整,比如480x360更快
    image_np_expanded = np.expand_dims(image_np, axis=0)
    
  • 简化可视化:把line_thickness改小(比如设为2),或者暂时关闭可视化,先测试推理速度,确认是可视化拖慢还是推理拖慢。

4. 调整waitKey的时间

你的代码里用了cv2.waitKey(25),对应40fps的刷新频率,但如果你的推理速度达不到这么快,这个时间可以适当调小,比如改成10ms,不过前面的优化才是重点。

按照这些步骤改完,延迟应该能降到1-2秒以内,卡顿也会明显减少。


内容的提问来源于stack exchange,提问作者Harry Boy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:19:42