目标检测模型训练:LabelMe多边形标注精度及XML转TFRecord方法咨询
Hey there! Let's break down your two questions step by step—they're both solid considerations for refining your object detection pipeline:
- Yes, in specific scenarios: If your detection targets are irregularly shaped (think curved roads, oddly manufactured parts, or natural objects like plants/animals), using LabelMe's polygon annotations will give you a far more precise outline of the target compared to bounding boxes. This cuts down on redundant background pixels that can confuse the model, especially if the target is surrounded by visually similar elements. I’ve seen mAP scores jump 3-5 points in these cases just by switching to polygon-based (or converted tight bounding box) annotations.
- Not so much for regular objects: If your targets are standard rectangular/square items (like books, boxes, or phones), polygon annotations won’t offer meaningful gains over well-placed bounding boxes—they’ll just add unnecessary annotation time and effort.
- A quick caveat: Make sure your model can handle polygon inputs, or plan to convert the polygons to tight bounding boxes first. Most classic detectors (Faster R-CNN, YOLO) expect bounding boxes, so you’ll need to compute the minimum enclosing rectangle for each polygon if you stick with these models.
LabelMe’s XML format stores polygon vertex coordinates instead of direct bounding boxes, so we need to parse those first, then package the data into TFRecord format. Here’s a practical, code-driven approach:
1. Parse LabelMe XML Files
First, we’ll extract key info (image details, class labels, polygon points) from the XMLs. We can use Python’s built-in xml.etree.ElementTree for this:
import xml.etree.ElementTree as ET def parse_labelme_xml(xml_path): tree = ET.parse(xml_path) root = tree.getroot() # Grab image metadata filename = root.find('filename').text size = root.find('size') width = int(size.find('width').text) height = int(size.find('height').text) objects = [] for obj in root.findall('object'): cls_name = obj.find('name').text # Extract polygon vertices polygon = obj.find('polygon') points = [(float(pt.find('x').text), float(pt.find('y').text)) for pt in polygon.findall('pt')] # Optional: Convert polygon to a tight bounding box (for standard detectors) x_min = min(p[0] for p in points) y_min = min(p[1] for p in points) x_max = max(p[0] for p in points) y_max = max(p[1] for p in points) objects.append({ 'class': cls_name, 'polygon': points, 'bbox': [x_min, y_min, x_max, y_max] }) return { 'filename': filename, 'width': width, 'height': height, 'objects': objects }
2. Convert Parsed Data to TFRecord
Next, we’ll wrap the parsed data into TensorFlow’s tf.train.Example format and write it to a TFRecord file. You’ll need a class-to-ID mapping for your dataset:
import tensorflow as tf import os def create_tf_example(data, image_dir, class_to_id): # Load the image as bytes image_path = os.path.join(image_dir, data['filename']) with tf.io.gfile.GFile(image_path, 'rb') as f: encoded_image = f.read() # Build the feature dictionary feature = { 'image/encoded': tf.train.Feature(bytes_list=tf.train.BytesList(value=[encoded_image])), 'image/filename': tf.train.Feature(bytes_list=tf.train.BytesList(value=[data['filename'].encode('utf-8')])), 'image/width': tf.train.Feature(int64_list=tf.train.Int64List(value=[data['width']])), 'image/height': tf.train.Feature(int64_list=tf.train.Int64List(value=[data['height']])), # Bounding box features (for standard detectors) 'image/object/bbox/xmin': tf.train.Feature(float_list=tf.train.FloatList(value=[obj['bbox'][0] for obj in data['objects']])), 'image/object/bbox/ymin': tf.train.Feature(float_list=tf.train.FloatList(value=[obj['bbox'][1] for obj in data['objects']])), 'image/object/bbox/xmax': tf.train.Feature(float_list=tf.train.FloatList(value=[obj['bbox'][2] for obj in data['objects']])), 'image/object/bbox/ymax': tf.train.Feature(float_list=tf.train.FloatList(value=[obj['bbox'][3] for obj in data['objects']])), 'image/object/class/label': tf.train.Feature(int64_list=tf.train.Int64List(value=[class_to_id[obj['class']] for obj in data['objects']])), 'image/object/class/text': tf.train.Feature(bytes_list=tf.train.BytesList(value=[obj['class'].encode('utf-8') for obj in data['objects']])), # Optional: Add polygon points if your model supports them # 'image/object/polygon/points': tf.train.Feature(float_list=tf.train.FloatList(value=[coord for obj in data['objects'] for pt in obj['polygon'] for coord in pt])) } return tf.train.Example(features=tf.train.Features(feature=feature)) def convert_to_tfrecord(xml_dir, image_dir, output_path, class_to_id): writer = tf.io.TFRecordWriter(output_path) for xml_file in os.listdir(xml_dir): if not xml_file.endswith('.xml'): continue xml_path = os.path.join(xml_dir, xml_file) data = parse_labelme_xml(xml_path) tf_example = create_tf_example(data, image_dir, class_to_id) writer.write(tf_example.SerializeToString()) writer.close() print(f"Successfully saved TFRecord to {output_path}")
3. Run the Conversion
Just plug in your paths and class mapping:
# Replace with your own class IDs and paths class_to_id = {'crop_field': 0, 'construction_vehicle': 1} xml_directory = '/path/to/your/labelme/xmls' image_directory = '/path/to/your/images' output_tfrecord = '/path/to/save/output.tfrecord' convert_to_tfrecord(xml_directory, image_directory, output_tfrecord, class_to_id)
内容的提问来源于stack exchange,提问作者akn

