如何基于手动分割真值准备Fast R-CNN训练数据以自动提取目标Bounding Box?
Absolutely, I’ve been through the process of prepping data for Fast R-CNN training multiple times, especially when working with manual segmentation masks as ground truth. Let me break down exactly how to adapt your existing images and masks into a format Fast R-CNN can use, plus share practical code snippets to streamline the work.
Fast R-CNN relies on bounding box (bbox) annotations, not segmentation masks, so your first task is to extract tight bboxes from your manual masks. For any object mask, the bbox is simply the smallest rectangle that encloses all non-zero pixels in the mask.
Here’s a Python snippet to automate this using OpenCV and NumPy:
import cv2 import numpy as np import os def mask_to_bbox(mask_path, class_label): # Load single-channel mask (adjust if your masks are multi-channel) mask = cv2.imread(mask_path, cv2.IMREAD_GRAYSCALE) # Find all non-zero pixel coordinates in the mask coords = np.column_stack(np.where(mask > 0)) if coords.size == 0: return None # Skip empty masks # Calculate bbox coordinates (x1, y1 = top-left; x2, y2 = bottom-right) y_min, x_min = coords.min(axis=0) y_max, x_max = coords.max(axis=0) return (x_min, y_min, x_max, y_max, class_label) # Configure paths and class mapping (adjust to your dataset) mask_dir = "/path/to/your/manual_masks" image_dir = "/path/to/your/test_images" class_mapping = {"cat": 1, "dog": 2} # Background = 0, target classes start at 1 # Generate bbox annotations for all masks bbox_annotations = [] for mask_file in os.listdir(mask_dir): if not mask_file.endswith(".png"): continue # Match mask to corresponding image (adjust naming logic to fit your files) img_file = mask_file.replace("_mask.png", ".jpg") img_path = os.path.join(image_dir, img_file) if not os.path.exists(img_path): print(f"Image {img_file} not found, skipping mask {mask_file}") continue # Extract class from mask filename (e.g., "cat_001_mask.png" → "cat") class_name = mask_file.split("_")[0] class_label = class_mapping.get(class_name, 0) if class_label == 0: print(f"Unknown class for mask {mask_file}, skipping") continue bbox = mask_to_bbox(os.path.join(mask_dir, mask_file), class_label) if bbox: bbox_annotations.append((img_path, *bbox))
Fast R-CNN (the pycaffe implementation) works best with either a PASCAL VOC-style directory structure or a simple text-based annotation format. Here’s how to set up both:
Option 1: PASCAL VOC-Style Structure (Recommended)
Create this directory layout:
your_custom_dataset/ ├── Annotations/ # XML files with bbox annotations ├── Images/ # All training/validation images ├── ImageSets/ │ └── Main/ │ ├── train.txt # Filenames (no extension) of training images │ └── val.txt # Filenames of validation images
Use this code to generate the required XML annotations and split files:
from xml.etree.ElementTree import Element, SubElement, tostring from xml.dom.minidom import parseString def create_voc_xml(img_path, bboxes, save_path): # Get image dimensions img = cv2.imread(img_path) height, width, depth = img.shape # Build XML structure annotation = Element('annotation') SubElement(annotation, 'folder').text = os.path.basename(os.path.dirname(img_path)) SubElement(annotation, 'filename').text = os.path.basename(img_path) SubElement(annotation, 'path').text = img_path size = SubElement(annotation, 'size') SubElement(size, 'width').text = str(width) SubElement(size, 'height').text = str(height) SubElement(size, 'depth').text = str(depth) SubElement(annotation, 'segmented').text = '0' # We're using bboxes, not segmentation # Add each bbox to the XML class_id_to_name = {v: k for k, v in class_mapping.items()} for bbox in bboxes: x1, y1, x2, y2, label = bbox obj = SubElement(annotation, 'object') SubElement(obj, 'name').text = class_id_to_name[label] SubElement(obj, 'pose').text = 'Unspecified' SubElement(obj, 'truncated').text = '0' SubElement(obj, 'difficult').text = '0' bndbox = SubElement(obj, 'bndbox') SubElement(bndbox, 'xmin').text = str(x1) SubElement(bndbox, 'ymin').text = str(y1) SubElement(bndbox, 'xmax').text = str(x2) SubElement(bndbox, 'ymax').text = str(y2) # Save formatted XML xml_str = parseString(tostring(annotation)).toprettyxml(indent=" ") with open(save_path, 'w') as f: f.write(xml_str) # Create annotations directory annotations_dir = "/path/to/your_custom_dataset/Annotations" os.makedirs(annotations_dir, exist_ok=True) # Group bboxes by image img_bbox_map = {} for item in bbox_annotations: img_path, x1, y1, x2, y2, label = item if img_path not in img_bbox_map: img_bbox_map[img_path] = [] img_bbox_map[img_path].append((x1, y1, x2, y2, label)) # Generate XML files for each image for img_path, bboxes in img_bbox_map.items(): xml_filename = os.path.basename(img_path).replace(".jpg", ".xml") xml_path = os.path.join(annotations_dir, xml_filename) create_voc_xml(img_path, bboxes, xml_path) # Split into train/val sets (80/20 split example) img_filenames = [os.path.splitext(os.path.basename(p))[0] for p in img_bbox_map.keys()] split_idx = int(len(img_filenames) * 0.8) train_files = img_filenames[:split_idx] val_files = img_filenames[split_idx:] # Save train/val split files with open("/path/to/your_custom_dataset/ImageSets/Main/train.txt", 'w') as f: f.write("\n".join(train_files)) with open("/path/to/your_custom_dataset/ImageSets/Main/val.txt", 'w') as f: f.write("\n".join(val_files))
Option 2: Simple Text Annotation Format
If you prefer to skip XMLs, create a text file where each line follows this format:image_path x1 y1 x2 y2 class_label
Use this code to generate the files:
# Save train/val annotations as text files train_annot_path = "/path/to/train_annotations.txt" val_annot_path = "/path/to/val_annotations.txt" # Split annotations (match the train/val split from earlier) train_items = [] val_items = [] train_img_set = set(train_files) for item in bbox_annotations: img_name = os.path.splitext(os.path.basename(item[0]))[0] if img_name in train_img_set: train_items.append(item) else: val_items.append(item) # Write to files with open(train_annot_path, 'w') as f: for item in train_items: f.write(f"{item[0]} {item[1]} {item[2]} {item[3]} {item[4]} {item[5]}\n") with open(val_annot_path, 'w') as f: for item in val_items: f.write(f"{item[0]} {item[1]} {item[2]} {item[3]} {item[4]} {item[5]}\n")
You’ll need to modify two core prototxt files to work with your dataset:
- Data Layer: Update the
annotated_data_paramto point to your dataset. For PASCAL VOC-style data:layer { name: "data" type: "AnnotatedData" top: "data" top: "im_info" top: "gt_boxes" include { phase: TRAIN } transform_param { mirror: true crop_size: 224 mean_value: 104 mean_value: 117 mean_value: 123 } annotated_data_param { batch_size: 2 # Adjust based on your GPU memory source: "/path/to/your_custom_dataset/ImageSets/Main/train.txt" anno_type: "PASCAL_VOC" root_folder: "/path/to/your_custom_dataset" } } - Classification Layer: Set
num_classesto your target class count + 1 (for the background class). For example, 2 target classes →num_classes: 3.
Use a pre-trained model from the Caffe Model Zoo (like VGG16) and run the training command:
python tools/train_net.py --gpu 0 --solver models/VGG16/fast_rcnn_solver.prototxt --weights data/imagenet_models/VGG16.v2.caffemodel
Adjust the --solver and --weights paths to match your local setup.
- Clean Masks: Incomplete or noisy segmentation masks will generate inaccurate bboxes, which hurt training quality. Double-check your masks before proceeding.
- Class Labels: Fast R-CNN reserves class 0 for the background, so always start your target class labels at 1.
- GPU Memory: If you hit out-of-memory errors, reduce the
batch_sizein the solver and data layer prototxts. - Augmentation: The transform param includes mirroring—add random cropping or brightness adjustments if you need better generalization.
内容的提问来源于stack exchange,提问作者sc241

