TensorFlow Object Detection检测阶段输出修改方法咨询
Hey there! Let me walk you through how to tweak the output of the TensorFlow Object Detection framework during inference, since you're working with Anaconda Python 3.6, TF 1.5, and camera input.
First, you'll need to target the script handling your detection logic:
- If you're using the official camera detection workflow, look for
object_detection/legacy/video.py(TF 1.5 relies on the legacy module for this use case). - If you've written a custom inference script for your camera feed, you'll modify that directly.
- For batch inference tasks, the core script is
object_detection/inference/infer_detections.py.
The standard workflow outputs bounding boxes by drawing them on frames. To customize this, you'll edit the section where detection results (detection_boxes, detection_scores, detection_classes) are processed.
2.1 Editing the Official video.py Script
Find the code block that handles drawing bounding boxes (it’ll look something like this for TF 1.5):
# Extract detection results boxes = np.squeeze(detection_boxes) scores = np.squeeze(detection_scores) classes = np.squeeze(detection_classes).astype(np.int32) # Default: Draw bounding boxes on frame for i in range(min(top_k, boxes.shape[0])): if scores[i] > score_thresh: # Calculate pixel coordinates from normalized values ymin, xmin, ymax, xmax = boxes[i] left = int(xmin * im_width) right = int(xmax * im_width) top = int(ymin * im_height) bottom = int(ymax * im_height) # Default drawing code cv2.rectangle(image_np, (left, top), (right, bottom), (0, 255, 0), 3) cv2.putText(image_np, f"{category_index[classes[i]]['name']}: {scores[i]:.2f}", (left, top-10), cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0,255,0), 2)
Here’s how to customize the output:
- Print to console: Add a print statement inside the loop to log results in real time:
print(f"Detected: {category_index[classes[i]]['name']} | Score: {scores[i]:.2f} | BBox: ({left}, {top}, {right}, {bottom})") - Save to a file: Use the
csvmodule to log results to a CSV for later analysis:import csv with open('camera_detections.csv', 'a', newline='') as f: writer = csv.writer(f) # Write header once at the start of the script # writer.writerow(["Class", "Score", "Left", "Top", "Right", "Bottom"]) writer.writerow([category_index[classes[i]]['name'], scores[i], left, top, right, bottom]) - Adjust output content: Filter results to only show high-confidence detections, add timestamp/frame number, or format outputs into JSON strings—all by modifying this loop.
2.2 Editing a Custom Inference Script
If you’re using a custom script, look for the section where you run the detection session:
with tf.Session() as sess: # Load your trained model ... # Process camera feed cap = cv2.VideoCapture(0) # 0 = default camera while True: ret, frame = cap.read() if not ret: break # Preprocess frame for model input image_np_expanded = np.expand_dims(frame, axis=0) # Run detection (boxes, scores, classes, num) = sess.run( [detection_boxes, detection_scores, detection_classes, num_detections], feed_dict={image_tensor: image_np_expanded}) # -------------------------- # Modify output HERE # -------------------------- boxes = np.squeeze(boxes) scores = np.squeeze(scores) classes = np.squeeze(classes).astype(np.int32) # Add your custom output logic here (same as the video.py examples)
- Since you’re on TF 1.5, stick to TF 1.x APIs (like
sess.run()instead of eager execution) to avoid compatibility issues. - Ensure your camera frame dimensions match the model’s input size—otherwise, bounding box coordinates will be misaligned.
- If you need to disable the default bounding box drawing entirely, just comment out the
cv2.rectangleandcv2.putTextlines.
内容的提问来源于stack exchange,提问作者Nolan H

