You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改YOLOv5实现实时检测后直接生成目标图像嵌入?

问题描述

我希望使用YOLOv5对摄像头画面进行实时目标检测,然后为检测到的目标裁剪图像生成向量嵌入。

目前我使用以下函数为本地保存的图片生成图像嵌入:

def generate_img_embedding(img_file_path):
    images = [
        Image.open(img_file_path)
    ]
    # Encoding a single image takes ~20 ms
    embeddings = embedding_model.encode(images)  # 修正原代码笔误:img_str改为images

    return embeddings

此外,我通过以下方式启动带图像裁剪功能的YOLOv5目标检测:

def start_camera(productid):
    print("Attempting to start camera")
    # productid = "11011"

    try:
        command = " python ./yolov5/detect.py  --source 0 --save-crop --name "+ productid +" --project ./cropped_images"  # 修正原代码变量错误:id改为productid
        os.system(command)
        print("Camera runnning")
    except Exception as e:
        print("error starting camera!", e)

请问如何修改YOLOv5模型,以实时将裁剪图像传入我的嵌入生成函数?


解决方案

命令行调用detect.py的方式无法实时传递裁剪图到你的嵌入函数,最优方案是直接在Python脚本中整合YOLOv5检测逻辑与向量生成流程,跳过本地文件读写环节,提升实时性。

步骤1:编写整合式实时检测脚本

放弃命令行调用,直接加载YOLOv5模型处理摄像头帧,实时裁剪目标并生成嵌入:

import cv2
from PIL import Image
import torch
from pathlib import Path

# 提前初始化你的嵌入模型(示例:如CLIP、自定义ResNet等)
# embedding_model = ... 

def generate_img_embedding(img):
    # 改为直接接收PIL图像,省去文件IO开销
    embeddings = embedding_model.encode([img])
    return embeddings

def realtime_detect_with_embedding(productid):
    print("Attempting to start camera")
    # 加载YOLOv5模型(本地加载或在线拉取,按需选择)
    # 本地加载:model = torch.hub.load('./yolov5', 'custom', path='./yolov5/yolov5s.pt', source='local')
    model = torch.hub.load('ultralytics/yolov5', 'yolov5s', pretrained=True)

    # 打开默认摄像头
    cap = cv2.VideoCapture(0)
    if not cap.isOpened():
        print("Failed to open camera")
        return

    # 可选:创建裁剪图保存目录(和原逻辑一致)
    save_dir = Path(f"./cropped_images/{productid}")
    save_dir.mkdir(parents=True, exist_ok=True)

    frame_count = 0
    print("Camera running")
    while True:
        ret, frame = cap.read()
        if not ret:
            print("Failed to grab frame")
            break

        # 格式转换:OpenCV BGR → PIL RGB
        img_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
        pil_img = Image.fromarray(img_rgb)

        # YOLOv5目标检测
        results = model(pil_img)

        # 遍历每一个检测结果
        for idx, det in enumerate(results.xyxy[0]):
            x1, y1, x2, y2, conf, cls_id = det
            # 裁剪目标区域
            crop_img = pil_img.crop((x1.item(), y1.item(), x2.item(), y2.item()))

            # 实时生成向量嵌入
            embedding = generate_img_embedding(crop_img)
            print(f"检测到目标:{model.names[int(cls_id)]},嵌入向量维度:{embedding.shape}")

            # 可选:保存裁剪图到本地
            crop_save_path = save_dir / f"frame_{frame_count}_obj_{idx}.jpg"
            crop_img.save(crop_save_path)

        # 显示检测画面(按q退出)
        cv2.imshow('YOLOv5实时检测', frame)
        if cv2.waitKey(1) & 0xFF == ord('q'):
            break

        frame_count += 1

    # 释放资源
    cap.release()
    cv2.destroyAllWindows()

# 启动检测
realtime_detect_with_embedding("11011")

步骤2:关键修改说明

  1. 直接加载YOLOv5模型:跳过命令行调用,直接在Python中获取检测结果的坐标数据,无需依赖本地文件传递。
  2. 优化嵌入生成函数:改为直接接收PIL图像,避免反复读写本地文件,大幅提升效率。
  3. 保留裁剪图保存功能:如果需要留存裁剪图,可按原逻辑创建目录并保存,无需依赖YOLOv5的--save-crop参数。

注意事项

  • 确保embedding_model已正确初始化,且支持PIL图像输入。
  • 若检测速度不足,可降低摄像头分辨率(如cap.set(cv2.CAP_PROP_FRAME_WIDTH, 640))或使用更轻量的YOLOv5模型(如yolov5n.pt)。

内容的提问来源于stack exchange,提问作者e.iluf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 13:45:23