You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将基于YOLOv3的单图像目标检测模型适配MP4视频?

如何将YOLOv3单图像检测修改为MP4视频检测?

原代码实现了单图像的YOLOv3目标检测,要适配MP4视频检测,核心是通过OpenCV的VideoCapture逐帧读取视频内容,重复执行单帧检测逻辑即可。以下是修改后的完整代码及关键说明:

关键修改点

  • 替换单图像读取逻辑为视频流读取:使用cv2.VideoCapture加载MP4文件
  • 新增循环帧处理逻辑:持续读取视频帧直到播放结束
  • 调整窗口交互逻辑:视频播放时需用waitKey(1)实现实时刷新,支持按键退出
  • 可选:添加视频写入逻辑,将检测结果保存为新的MP4文件

修改后的完整代码

import cv2
import numpy as np

# Load Yolo
net = cv2.dnn.readNet("yolov3.weights", "yolov3.cfg")
classes = []
with open("coco.names", "r") as f:
    classes = [line.strip() for line in f.readlines()]
layer_names = net.getLayerNames()
output_layers = [layer_names[i - 1] for i in net.getUnconnectedOutLayers()]
colors = np.random.uniform(0, 255, size=(len(classes), 3))

# 加载MP4视频,替换为你的视频路径
cap = cv2.VideoCapture("input.mp4")

# 可选:初始化视频写入器,保存检测结果
fourcc = cv2.VideoWriter_fourcc(*'mp4v')
frame_width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
frame_height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
fps = cap.get(cv2.CAP_PROP_FPS)
out = cv2.VideoWriter("output.mp4", fourcc, fps, (frame_width, frame_height))

while cap.isOpened():
    ret, frame = cap.read()
    if not ret:
        break  # 视频播放结束
    
    height, width, channels = frame.shape
    blob = cv2.dnn.blobFromImage(frame, 0.00392, (416, 416), (0, 0, 0), True, crop=False)

    net.setInput(blob)
    outs = net.forward(output_layers)

    class_ids = []
    confidences = []
    boxes = []
    for out in outs:
        for detection in out:
            scores = detection[5:]
            class_id = np.argmax(scores)
            confidence = scores[class_id]
            if confidence > 0.5:
                # 计算目标框坐标
                center_x = int(detection[0] * width)
                center_y = int(detection[1] * height)
                w = int(detection[2] * width)
                h = int(detection[3] * height)

                x = int(center_x - w / 2)
                y = int(center_y - h / 2)
                
                boxes.append([x, y, w, h])
                confidences.append(float(confidence))
                class_ids.append(class_id)
    
    # 非极大值抑制去除重复框
    indexes = cv2.dnn.NMSBoxes(boxes, confidences, 0.5, 0.4)
    font = cv2.FONT_HERSHEY_PLAIN
    for i in range(len(boxes)):
        if i in indexes:   
            x, y, w, h = boxes[i]
            label = str(classes[class_ids[i]])
            color = colors[class_ids[i]]  # 改为按类别固定颜色,避免帧间颜色变化
            cv2.rectangle(frame, (x, y), (x + w, y + h), color, 2)
            cv2.putText(frame, label, (x, y + 30), font, 3, color, 3)
    
    # 显示当前帧
    cv2.imshow("Video Detection", frame)
    # 写入输出视频(可选)
    out.write(frame)
    
    # 按下q键退出播放
    if cv2.waitKey(1) & 0xFF == ord('q'):
        break

# 释放资源
cap.release()
out.release()
cv2.destroyAllWindows()

额外说明

  1. 确保yolov3.weights、yolov3.cfg、coco.names文件路径正确
  2. 替换input.mp4为你的目标视频路径,output.mp4为保存结果的路径
  3. 代码中把原有的colors[i]改为colors[class_ids[i]],可以让同一类别的目标在所有帧中保持相同颜色,提升视觉一致性
  4. 如果不需要保存输出视频,可以删除视频写入相关的代码块

内容的提问来源于stack exchange,提问作者Kavishka Rajapakshe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 14:05:20