You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将TensorFlow目标检测API的检测框坐标保存为CSV文件?

Got it, let's walk through how to save those detection box coordinates to a CSV file. Here's how you can modify your existing TensorFlow Object Detection API code to make this happen:

Step 1: Add Required Imports

First, you'll need the csv module to handle writing the CSV file, and os to extract image filenames easily:

import csv
import os

Step 2: Full Modified Code

Integrate the CSV writing logic directly into your existing detection loop. The key parts are converting normalized box coordinates to actual pixel values, filtering low-confidence detections, and writing each valid result to the CSV:

# Define your CSV output filename
csv_output = 'detection_boxes.csv'

# Initialize the CSV file with headers
with open(csv_output, 'w', newline='') as csv_file:
    # Define the columns we want to save
    fieldnames = [
        'image_name', 
        'ymin_pixel', 'xmin_pixel', 
        'ymax_pixel', 'xmax_pixel',
        'confidence_score', 
        'class_id', 'class_name'
    ]
    writer = csv.DictWriter(csv_file, fieldnames=fieldnames)
    writer.writeheader()

    with detection_graph.as_default():
        with tf.Session(graph=detection_graph) as sess:
            # Get model input/output tensors (your existing code)
            image_tensor = detection_graph.get_tensor_by_name('image_tensor:0')
            detection_boxes = detection_graph.get_tensor_by_name('detection_boxes:0')
            detection_scores = detection_graph.get_tensor_by_name('detection_scores:0')
            detection_classes = detection_graph.get_tensor_by_name('detection_classes:0')
            num_detections = detection_graph.get_tensor_by_name('num_detections:0')

            for image_path in TEST_IMAGE_PATHS:
                image = Image.open(image_path)
                im_width, im_height = image.size  # Get actual image dimensions
                image_np = load_image_into_numpy_array(image)
                image_np_expanded = np.expand_dims(image_np, axis=0)

                # Run detection (your existing code)
                (boxes, scores, classes, num) = sess.run(
                    [detection_boxes, detection_scores, detection_classes, num_detections],
                    feed_dict={image_tensor: image_np_expanded})

                # Visualize results (keep your existing visualization code)
                vis_util.visualize_boxes_and_labels_on_image_array(
                    image_np,
                    np.squeeze(boxes),
                    np.squeeze(classes).astype(np.int32),
                    np.squeeze(scores),
                    category_index,
                    use_normalized_coordinates=True,
                    line_thickness=8)
                plt.figure(figsize=IMAGE_SIZE)
                plt.imshow(image_np)

                # Process and save detection boxes to CSV
                boxes_squeezed = np.squeeze(boxes)
                scores_squeezed = np.squeeze(scores)
                classes_squeezed = np.squeeze(classes).astype(np.int32)
                image_name = os.path.basename(image_path)  # Get just the filename, not full path

                # Loop through each detection (filter low-confidence results)
                for i in range(int(num[0])):
                    confidence = scores_squeezed[i]
                    # Skip detections with confidence below your threshold (adjust as needed)
                    if confidence < 0.5:
                        continue

                    # Convert normalized coordinates to actual pixel values
                    # Model outputs boxes in [ymin, xmin, ymax, xmax] format (normalized 0-1)
                    ymin = boxes_squeezed[i][0] * im_height
                    xmin = boxes_squeezed[i][1] * im_width
                    ymax = boxes_squeezed[i][2] * im_height
                    xmax = boxes_squeezed[i][3] * im_width

                    # Get class info
                    class_id = classes_squeezed[i]
                    class_name = category_index[class_id]['name']

                    # Write the result to CSV
                    writer.writerow({
                        'image_name': image_name,
                        'ymin_pixel': round(ymin, 2),
                        'xmin_pixel': round(xmin, 2),
                        'ymax_pixel': round(ymax, 2),
                        'xmax_pixel': round(xmax, 2),
                        'confidence_score': round(confidence, 4),
                        'class_id': class_id,
                        'class_name': class_name
                    })

Key Details Explained:

  • Coordinate Conversion: The model returns normalized box coordinates (0 to 1) relative to the image size. Multiply by im_height and im_width to get actual pixel positions.
  • Confidence Filter: We skip detections with confidence < 0.5 to avoid cluttering the CSV with low-quality results—adjust this threshold based on your needs.
  • CSV Structure: Each row in the CSV corresponds to one valid detection box, with all relevant metadata (image name, coordinates, confidence, class info) for easy post-processing.

Once you run the code, you'll get a detection_boxes.csv file in your working directory with all the detection box data you need.

内容的提问来源于stack exchange,提问作者Ajinkya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:50:18