YOLOv8视频帧检测异常求助:non-inheat帧误堆叠至inheat列表
问题场景与异常
本人毕业论文需使用包含inheat和non-inheat两类的YOLOv8模型对视频帧进行检测分类。目前代码运行出现异常:模型可识别出non-inheat类别,但对应帧却被错误添加到inheat帧列表,而非non-inheat列表。已查阅YOLOv8官方文档未解决,使用PyCharm开发,代码如下:
def process_video_with_second_model(video_path): cap = cv2.VideoCapture(video_path) class_counts = {'inheat': 0, 'non-inheat': 0} in_heat_frames = [] non_in_heat_frames = [] while True: ret, frame = cap.read() if frame is None: break # Break the loop when no more frames are available # Resize the frame to a smaller size (e.g., 400x400) frame_small = cv2.resize(frame, (400, 400)) # Use the second model to detect in-heat behavior results_in_heat = yolov8_model_in_heat.predict(source=frame_small, show=True, conf=0.8) # Print results to inspect structure for results_in_heat_instance in results_in_heat: # Access bounding box coordinates boxes = results_in_heat_instance.boxes # CONFIDENCE 0.5 if len(boxes) > 0: class_name = results_in_heat_instance.names[0] # Use a dictionary to store the counts for each class class_counts[class_name] += 1 # Add the frame to the corresponding list based on the class name if class_name == 'non-inheat': non_in_heat_frames.append(frame) elif class_name == 'inheat': in_heat_frames.append(frame) print(f"Class Counts: {class_counts}") # Check if either condition is met (50 frames for inheat and 50 frames for non-inheat) if class_counts['inheat'] >= 50 and class_counts['non-inheat'] >= 50: break # Release resources for the second model cap.release() cv2.destroyAllWindows() # Stack the in-heat and non-in-heat frames vertically stacked_in_heat_frames = np.vstack(in_heat_frames) stacked_non_in_heat_frames = np.vstack(non_in_heat_frames) # Display the stacked in-heat and non-in-heat frames cv2.imshow('Stacked In-Heat Frames', stacked_in_heat_frames) cv2.imshow('Stacked Non-In-Heat Frames', stacked_non_in_heat_frames) cv2.waitKey(0) cv2.destroyAllWindows() # Compare the counts and return the label with the higher count if class_counts['inheat'] > class_counts['non-inheat']: return 'inheat' elif class_counts['non-inheat'] > class_counts['inheat']: return 'non-inheat'
问题根源
核心错误在于类别获取逻辑:
- YOLOv8的
results.names是类别ID到类名的映射字典,而非检测结果的排序列表。代码中直接取names[0],会固定获取字典的第一个键值(通常是训练时的第一个类别,比如inheat),导致无论检测到哪类,都被判定为inheat。 - 额外问题:同一帧若存在多个检测框,会重复计数并多次添加同一帧到列表,不符合单帧单分类的需求。
修复后的代码
def process_video_with_second_model(video_path): cap = cv2.VideoCapture(video_path) class_counts = {'inheat': 0, 'non-inheat': 0} in_heat_frames = [] non_in_heat_frames = [] while True: ret, frame = cap.read() if not ret or frame is None: break # 无更多帧时退出循环 # 缩小帧尺寸 frame_small = cv2.resize(frame, (400, 400)) # 模型检测 results_in_heat = yolov8_model_in_heat.predict(source=frame_small, show=True, conf=0.8) current_frame_class = None for result in results_in_heat: boxes = result.boxes if len(boxes) > 0: # 取置信度最高的检测框对应的类别(避免单帧多框重复统计) max_conf_idx = boxes.conf.argmax() class_id = boxes.cls[max_conf_idx].item() current_frame_class = result.names[class_id] # 单帧仅统计一次 if current_frame_class is not None: class_counts[current_frame_class] += 1 if current_frame_class == 'non-inheat': non_in_heat_frames.append(frame) elif current_frame_class == 'inheat': in_heat_frames.append(frame) print(f"类别计数: {class_counts}") # 收集到足够样本后停止 if class_counts['inheat'] >= 50 and class_counts['non-inheat'] >= 50: break # 释放资源 cap.release() cv2.destroyAllWindows() # 统一帧尺寸后再堆叠(避免尺寸不一致报错) if in_heat_frames: resized_in_heat = [cv2.resize(f, (400, 400)) for f in in_heat_frames] stacked_in_heat = np.vstack(resized_in_heat) cv2.imshow('堆叠In-Heat帧', stacked_in_heat) if non_in_heat_frames: resized_non_in_heat = [cv2.resize(f, (400, 400)) for f in non_in_heat_frames] stacked_non_in_heat = np.vstack(resized_non_in_heat) cv2.imshow('堆叠Non-In-Heat帧', stacked_non_in_heat) cv2.waitKey(0) cv2.destroyAllWindows() # 返回计数更高的类别 if class_counts['inheat'] > class_counts['non-inheat']: return 'inheat' elif class_counts['non-inheat'] > class_counts['inheat']: return 'non-inheat' else: return 'tie' # 处理平局情况
关键修复点
- 精准获取检测类别:从置信度最高的检测框中提取类别ID,再通过
names映射到对应类名,确保类别判断准确。 - 单帧单次统计:一帧仅根据最高置信度结果做一次分类,避免重复计数和重复添加帧。
- 帧尺寸统一:堆叠前统一所有帧的尺寸,解决
np.vstack因尺寸不一致报错的问题。 - 严谨的视频读取判断:用
not ret or frame is None判断视频读取结束,逻辑更严谨。
内容的提问来源于stack exchange,提问作者Thesis Project
相关产品推荐
相关产品推荐

