TensorRT部署YOLOv5后推理结果解析异常求助
YOLOv5 TensorRT推理后处理问题解决方案
先解答输出维度85的疑问
YOLOv5的每个检测框输出包含85个元素,对应:
- 前4个:边界框的中心x、中心y、宽度w、高度h(均为相对于输入图像的归一化值)
- 第5个:目标存在置信度(obj_conf),表示该框包含目标的概率
- 后80个:80个类别的分类置信度(cls_conf),每个值对应框属于对应类别的概率
最终有效检测分数是obj_conf * cls_conf,而非单独用分类置信度判断。
核心问题分析
你的代码存在张量索引计算完全错误,导致取到的不是预期数据,进而出现置信度异常值、框数据为0的情况:
- 边界框索引错误:你用
cpu_output[4*i]获取框数据,但每个框的85个值是连续存储的,第i个框的起始索引应为i * n_classes(即i*85) - 置信度计算错误:直接取
cpu_output[i * n_classes + j]且j从1开始,既没取到正确的分类置信度,也没有结合目标置信度计算最终分数 - 未做非极大值抑制(NMS),会保留大量重复框
修正后的完整代码
void postprocessAndDisplay(cv::Mat &img, float *gpu_output, const Dims dims, float threshold){ // 拷贝GPU输出到CPU size_t dimsSize = accumulate(dims.d+1, dims.d+dims.nbDims, 1, multiplies<size_t>()); vector<float> cpu_output(dimsSize); cudaMemcpy(cpu_output.data(), gpu_output, cpu_output.size()*sizeof(float), cudaMemcpyDeviceToHost); vector<int> classIds; vector<cv::Rect> boxes; vector<float> confidences; int img_width = img.cols; int img_height = img.rows; int n_boxes = dims.d[1], n_classes = dims.d[2]; for (int i = 0; i < n_boxes; i++){ // 当前框的起始索引 int box_start_idx = i * n_classes; // 1. 获取目标存在置信度,低于阈值直接跳过 float obj_conf = cpu_output[box_start_idx + 4]; if (obj_conf < threshold) continue; // 2. 找到分类置信度最高的类别 uint32_t maxClass = 0; float max_cls_conf = 0.0f; for (int j = 0; j < 80; j++){ float cls_conf = cpu_output[box_start_idx + 5 + j]; if (cls_conf > max_cls_conf){ max_cls_conf = cls_conf; maxClass = j; } } // 3. 计算最终检测分数,低于阈值跳过 float final_score = obj_conf * max_cls_conf; if (final_score < threshold) continue; // 4. 解析边界框并转换为图像实际尺寸 float x = cpu_output[box_start_idx + 0]; float y = cpu_output[box_start_idx + 1]; float w = cpu_output[box_start_idx + 2]; float h = cpu_output[box_start_idx + 3]; int left = static_cast<int>((x - w / 2) * img_width); int top = static_cast<int>((y - h / 2) * img_height); int width = static_cast<int>(w * img_width); int height = static_cast<int>(h * img_height); // 确保框在图像范围内 left = max(0, left); top = max(0, top); width = min(img_width - left, width); height = min(img_height - top, height); // 保存结果 boxes.push_back(cv::Rect(left, top, width, height)); classIds.push_back(maxClass); confidences.push_back(final_score); } // 5. 执行NMS去除重复框 vector<int> indices; cv::dnn::NMSBoxes(boxes, confidences, threshold, 0.5f, indices); // 6. 绘制检测结果 for (int idx : indices){ cv::Rect box = boxes[idx]; int classId = classIds[idx]; float score = confidences[idx]; cv::rectangle(img, box, cv::Scalar(0, 255, 0), 2); string label = cv::format("Class %d: %.2f", classId, score); cv::putText(img, label, cv::Point(box.x, box.y - 10), cv::FONT_HERSHEY_SIMPLEX, 0.5, cv::Scalar(0, 255, 0), 2); } cv::resize(img, img, cv::Size(1000, 1000)); cv::imshow("Test", img); cv::waitKey(0); }
额外注意事项
- 如果导出的.engine模型未集成sigmoid层,所有输出值都是原始logits,需要手动转换:
float val = 1.0f / (1.0f + exp(-raw_val));,否则会出现超过1的异常值 - 确保推理输入尺寸和模型训练/导出时的尺寸一致(如640x640),否则边界框转换会出错
- NMS的IoU阈值(代码中0.5f)可按需调整,值越小保留的框越少,过滤越严格
内容的提问来源于stack exchange,提问作者Gronnmann
相关产品推荐
相关产品推荐

