如何使用TensorFlow Object Detection API提取bounding box内的图像?
提取检测框内图像的解决方案
你的代码已经实现了目标检测和可视化,但缺少提取检测框区域图像的逻辑。以下是修改后的完整代码,加入了裁剪并保存检测框内图像的功能:
while True: # 修正:cap是VideoCapture对象,应该用read()获取帧,而非imread() ret, frame = cap.read() if not ret: break # Expand dimensions since the model expects images to have shape: [1, None, None, 3] expanded_frame = np.expand_dims(frame, axis=0) (boxes, scores, classes, num_c) = sess.run( [detection_boxes, detection_scores, detection_classes, num_detections], feed_dict={image_tensor: expanded_frame}) # 获取图像尺寸,用于转换归一化坐标 h, w, _ = frame.shape # 去除冗余维度 boxes = np.squeeze(boxes) scores = np.squeeze(scores) classes = np.squeeze(classes).astype(np.int32) # 遍历检测结果,提取符合分数阈值的框内图像 for idx in range(len(scores)): if scores[idx] >= 0.40: # 转换归一化坐标为实际像素坐标(注意boxes格式是[ymin, xmin, ymax, xmax]) ymin, xmin, ymax, xmax = boxes[idx] x1 = int(xmin * w) y1 = int(ymin * h) x2 = int(xmax * w) y2 = int(ymax * h) # 裁剪检测框区域 cropped_image = frame[y1:y2, x1:x2] # 保存裁剪后的图像,用索引区分避免覆盖 save_path = fr'C:\Users\user\Desktop\New folder\Face Detected\cropped_{idx}.png' cv2.imwrite(save_path, cropped_image) # Visualization of the detection vis_util.visualize_boxes_and_labels_on_image_array( frame, np.squeeze(boxes), classes, scores, category_index, use_normalized_coordinates=True, line_thickness=1, min_score_thresh=0.40) cv2.imshow('Detection', frame) filename = r'C:\Users\user\Desktop\New folder\Face Detected\Detected.png' cv2.imwrite(filename, frame) fps.update() if cv2.waitKey(1) == ord('q'): fps.stop() break print("Fps: {:.2f}".format(fps.fps())) cap.release() # 修正:用release()而非stop(),VideoCapture对象没有stop()方法 cv2.destroyAllWindows()
关键说明:
- 坐标转换:TensorFlow Detection API返回的
boxes是归一化坐标(范围0-1),需要乘以图像的实际宽高转换为像素坐标,注意坐标顺序是[ymin, xmin, ymax, xmax] - 筛选有效框:只处理分数大于
min_score_thresh(这里是0.40)的检测结果,避免无效框 - 避免文件覆盖:保存裁剪图像时加入索引
idx,防止每次循环的图像互相覆盖 - 代码修正:原代码中
frame = cv2.imread(cap)是错误的,cap作为视频捕获对象应该用cap.read()获取帧;cap.stop()改为cap.release(),符合OpenCV的API规范
内容的提问来源于stack exchange,提问作者Pooja
相关产品推荐
相关产品推荐

