You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

YOLOv8视频帧检测异常求助:non-inheat帧误堆叠至inheat列表

问题场景与异常

本人毕业论文需使用包含inheat和non-inheat两类的YOLOv8模型对视频帧进行检测分类。目前代码运行出现异常:模型可识别出non-inheat类别,但对应帧却被错误添加到inheat帧列表,而非non-inheat列表。已查阅YOLOv8官方文档未解决,使用PyCharm开发,代码如下:

def process_video_with_second_model(video_path):
    cap = cv2.VideoCapture(video_path)
    class_counts = {'inheat': 0, 'non-inheat': 0}

    in_heat_frames = []
    non_in_heat_frames = []

    while True:
        ret, frame = cap.read()
        if frame is None:
            break # Break the loop when no more frames are available

        # Resize the frame to a smaller size (e.g., 400x400)
        frame_small = cv2.resize(frame, (400, 400))

        # Use the second model to detect in-heat behavior
        results_in_heat = yolov8_model_in_heat.predict(source=frame_small, show=True, conf=0.8)

        # Print results to inspect structure
        for results_in_heat_instance in results_in_heat:
            # Access bounding box coordinates
            boxes = results_in_heat_instance.boxes

            # CONFIDENCE 0.5
            if len(boxes) > 0:
                class_name = results_in_heat_instance.names[0]

                # Use a dictionary to store the counts for each class
                class_counts[class_name] += 1

                # Add the frame to the corresponding list based on the class name
                if class_name == 'non-inheat':
                    non_in_heat_frames.append(frame)
                elif class_name == 'inheat':
                    in_heat_frames.append(frame)

        print(f"Class Counts: {class_counts}")

        # Check if either condition is met (50 frames for inheat and 50 frames for non-inheat)
        if class_counts['inheat'] >= 50 and class_counts['non-inheat'] >= 50:
            break

    # Release resources for the second model
    cap.release()
    cv2.destroyAllWindows()

    # Stack the in-heat and non-in-heat frames vertically
    stacked_in_heat_frames = np.vstack(in_heat_frames)
    stacked_non_in_heat_frames = np.vstack(non_in_heat_frames)

    # Display the stacked in-heat and non-in-heat frames
    cv2.imshow('Stacked In-Heat Frames', stacked_in_heat_frames)
    cv2.imshow('Stacked Non-In-Heat Frames', stacked_non_in_heat_frames)
    cv2.waitKey(0)
    cv2.destroyAllWindows()

    # Compare the counts and return the label with the higher count
    if class_counts['inheat'] > class_counts['non-inheat']:
        return 'inheat'
    elif class_counts['non-inheat'] > class_counts['inheat']:
        return 'non-inheat'

问题根源

核心错误在于类别获取逻辑:

  • YOLOv8的results.names是类别ID到类名的映射字典,而非检测结果的排序列表。代码中直接取names[0],会固定获取字典的第一个键值(通常是训练时的第一个类别,比如inheat),导致无论检测到哪类,都被判定为inheat。
  • 额外问题:同一帧若存在多个检测框,会重复计数并多次添加同一帧到列表,不符合单帧单分类的需求。

修复后的代码

def process_video_with_second_model(video_path):
    cap = cv2.VideoCapture(video_path)
    class_counts = {'inheat': 0, 'non-inheat': 0}

    in_heat_frames = []
    non_in_heat_frames = []

    while True:
        ret, frame = cap.read()
        if not ret or frame is None:
            break  # 无更多帧时退出循环

        # 缩小帧尺寸
        frame_small = cv2.resize(frame, (400, 400))

        # 模型检测
        results_in_heat = yolov8_model_in_heat.predict(source=frame_small, show=True, conf=0.8)

        current_frame_class = None
        for result in results_in_heat:
            boxes = result.boxes
            if len(boxes) > 0:
                # 取置信度最高的检测框对应的类别(避免单帧多框重复统计)
                max_conf_idx = boxes.conf.argmax()
                class_id = boxes.cls[max_conf_idx].item()
                current_frame_class = result.names[class_id]

        # 单帧仅统计一次
        if current_frame_class is not None:
            class_counts[current_frame_class] += 1
            if current_frame_class == 'non-inheat':
                non_in_heat_frames.append(frame)
            elif current_frame_class == 'inheat':
                in_heat_frames.append(frame)

        print(f"类别计数: {class_counts}")

        # 收集到足够样本后停止
        if class_counts['inheat'] >= 50 and class_counts['non-inheat'] >= 50:
            break

    # 释放资源
    cap.release()
    cv2.destroyAllWindows()

    # 统一帧尺寸后再堆叠(避免尺寸不一致报错)
    if in_heat_frames:
        resized_in_heat = [cv2.resize(f, (400, 400)) for f in in_heat_frames]
        stacked_in_heat = np.vstack(resized_in_heat)
        cv2.imshow('堆叠In-Heat帧', stacked_in_heat)
    if non_in_heat_frames:
        resized_non_in_heat = [cv2.resize(f, (400, 400)) for f in non_in_heat_frames]
        stacked_non_in_heat = np.vstack(resized_non_in_heat)
        cv2.imshow('堆叠Non-In-Heat帧', stacked_non_in_heat)
    
    cv2.waitKey(0)
    cv2.destroyAllWindows()

    # 返回计数更高的类别
    if class_counts['inheat'] > class_counts['non-inheat']:
        return 'inheat'
    elif class_counts['non-inheat'] > class_counts['inheat']:
        return 'non-inheat'
    else:
        return 'tie'  # 处理平局情况

关键修复点

  1. 精准获取检测类别:从置信度最高的检测框中提取类别ID,再通过names映射到对应类名,确保类别判断准确。
  2. 单帧单次统计:一帧仅根据最高置信度结果做一次分类,避免重复计数和重复添加帧。
  3. 帧尺寸统一:堆叠前统一所有帧的尺寸,解决np.vstack因尺寸不一致报错的问题。
  4. 严谨的视频读取判断:用not ret or frame is None判断视频读取结束,逻辑更严谨。

内容的提问来源于stack exchange,提问作者Thesis Project

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 16:10:37