You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

YOLOv3多类别目标检测实现中所有类别置信度为0的技术问题排查请求

问题排查与解决方法

我仔细梳理了你的代码和遇到的问题,发现几个核心错误直接导致了类别概率全0的异常,下面逐一说明并给出修复方案:

1. 类别置信度提取逻辑错误

在findObjects函数中,你写了:

confs=scores[classIds]

这里的classIds是一个列表,而当前循环的classId才是当前检测框对应的类别索引(单个整数)。用列表去索引数组会导致错误的取值,甚至直接拿到全0的结果。

修复后:

confs = scores[classId]

2. NMS函数参数错误

调用cv2.dnn.NMSBoxes时,你传入了confs(单个值)作为置信度列表,但NMS需要的是你收集的所有检测框的置信度列表confidence:

indices=cv2.dnn.NMSBoxes(bound_box,confs,Confidence_Threshould,NMS_Threshould)

修复后:

indices=cv2.dnn.NMSBoxes(bound_box, confidence, Confidence_Threshould, NMS_Threshould)

3. OpenCV文本绘制函数拼写错误

代码中多次使用cv2.puttext,但正确的函数名是cv2.putText(注意大写的T),这个错误会导致无法绘制类别标签和FPS信息,虽然不直接导致类别全0,但会影响最终可视化效果。

修复后示例:

cv2.putText(image, f'{classes[classIds[i]]}{int(confidence[i]*100)}%', (x,y-10), cv2.FONT_HERSHEY_PLAIN, 0.6, (0,255,0), 2)

额外检查点

除了代码错误,你还可以确认以下几点:

  • 确保coco.names文件路径正确,文件内容完整(包含80个COCO类别名称);
  • 预训练权重文件(yolov3.weights/yolov3-tiny.weights)下载完整,没有损坏;
  • 视频输入的帧读取正常,success变量为True时才进行后续处理(当前代码的try-except可以优化,避免跳过有效帧)。

修正后的findObjects函数完整代码

def findObjects(outputs, image):
    h, w, c = image.shape
    bound_box = []
    classIds = []
    confidence = []
    for output in outputs:
        for detection in output:
            scores = detection[5:]
            classId = np.argmax(scores)
            confs = scores[classId]  # 修复:用单个classId索引
            if confs > Confidence_Threshould:
                # 优化:用原始图像的宽高计算框坐标,而非固定的320
                box_w = int(detection[2] * w)
                box_h = int(detection[3] * h)
                x = int((detection[0] * w) - (box_w / 2))
                y = int((detection[1] * h) - (box_h / 2))
                bound_box.append([x, y, box_w, box_h])
                classIds.append(classId)
                confidence.append(float(confs))
    print(len(bound_box))
    # 修复:传入confidence列表
    indices = cv2.dnn.NMSBoxes(bound_box, confidence, Confidence_Threshould, NMS_Threshould)
    # 兼容不同OpenCV版本的indices返回格式
    if len(indices) > 0:
        for i in indices.flatten():
            box = bound_box[i]
            x, y, box_w, box_h = box[0], box[1], box[2], box[3]
            cv2.rectangle(image, (x, y), (x + box_w, y + box_h), (255, 0, 0), 2)
            # 修复putText拼写并优化文本显示
            cv2.putText(image, f'{classes[classIds[i]]} {int(confidence[i]*100)}%', 
                        (x, y-10), cv2.FONT_HERSHEY_PLAIN, 0.8, (0,255,0), 2)
            cv2.putText(image, f'FPS:{fps:.1f}', (10, 30), cv2.FONT_HERSHEY_PLAIN, 1, (0,255,0), 2)
            # 优化时间戳显示:获取当前帧的实时时间戳
            current_timestamp = cap.get(cv2.CAP_PROP_POS_MSEC)
            cv2.putText(image, f'Timestamp:{current_timestamp:.0f}ms', (10, 60), cv2.FONT_HERSHEY_PLAIN, 1, (0,255,0), 2)

另外补充一个小优化:之前计算 bounding box 时用了固定的Width=320,但应该用原始图像的宽高w和h,这样框的位置才会准确对应到原视频帧上。

内容的提问来源于stack exchange,提问作者CoderWithDIfficulties

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 12:02:34