如何将MatLab.mat格式的姿态数据集转为YOLO训练所需文本格式?
Great question! Converting MATLAB .mat ground truth labels to YOLO's text format is totally doable with Python—here's a step-by-step guide tailored to the FLIC and MPI-INF human pose datasets you're working with:
Since you're training YOLO to detect body joints, we'll treat each joint as a small "object" for detection. The standard YOLO text label format requires one line per joint, with values normalized to the image's width/height (0-1 range):
class_id x_center y_center width height
If you're using YOLOv8/v9's dedicated keypoint mode, the format shifts to include raw joint coordinates instead of boxes—we'll note how to adjust for that later.
First, install the libraries needed to read .mat files and process images:
pip install scipy pillow numpy opencv-python
(Note: Use h5py instead of scipy if you're working with MATLAB 7.3+ .mat files, as they use a different storage format.)
The FLIC dataset's .mat files store image paths and joint coordinates in a straightforward structure. Here's a script to convert them:
import scipy.io as sio import numpy as np from PIL import Image import os # Configure paths and parameters MAT_FILE = "path/to/your/examples.mat" IMAGE_DIR = "path/to/flic/image/directory" LABEL_OUTPUT_DIR = "path/to/save/yolo/labels" # Map FLIC's 6 joints to class IDs (adjust based on your training needs) JOINT_CLASS_MAP = {0: 0, 1: 1, 2: 2, 3: 3, 4: 4, 5: 5} # Fixed box size for joints (adjust based on your image resolution) JOINT_BOX_PIXELS = 20 # Create output directory if it doesn't exist os.makedirs(LABEL_OUTPUT_DIR, exist_ok=True) # Load the MATLAB data mat_data = sio.loadmat(MAT_FILE) img_paths = mat_data['imgpaths'][:, 0] joint_coords = mat_data['coords'] # Shape: (6 joints, 2 coords, N samples) # Process each sample for sample_idx in range(len(img_paths)): # Get image info img_full_path = img_paths[sample_idx][0] img_name = os.path.basename(img_full_path) img_path = os.path.join(IMAGE_DIR, img_name) # Get image dimensions for normalization with Image.open(img_path) as img: img_w, img_h = img.size # Create corresponding label file label_file = os.path.join(LABEL_OUTPUT_DIR, f"{os.path.splitext(img_name)[0]}.txt") with open(label_file, 'w') as f: # Iterate over each joint for joint_idx in range(joint_coords.shape[0]): x, y = joint_coords[joint_idx, :, sample_idx] # Skip unannotated joints (marked with 0,0) if x == 0 and y == 0: continue # Normalize values to YOLO format x_center = x / img_w y_center = y / img_h box_w = JOINT_BOX_PIXELS / img_w box_h = JOINT_BOX_PIXELS / img_h # Write to label file f.write(f"{JOINT_CLASS_MAP[joint_idx]} {x_center:.6f} {y_center:.6f} {box_w:.6f} {box_h:.6f}\n")
The MPI-INF dataset uses a more nested .mat structure. Here's an adapted script for it:
import scipy.io as sio import numpy as np from PIL import Image import os # Configure paths and parameters MAT_FILE = "path/to/your/mpi_train.mat" IMAGE_DIR = "path/to/mpi/image/directory" LABEL_OUTPUT_DIR = "path/to/save/yolo/labels" # Map MPI's 16 joints to class IDs JOINT_CLASS_MAP = {i: i for i in range(16)} JOINT_BOX_PIXELS = 20 os.makedirs(LABEL_OUTPUT_DIR, exist_ok=True) # Load MATLAB data (struct_as_record preserves nested structure) mat_data = sio.loadmat(MAT_FILE, struct_as_record=False) anno_list = mat_data['annolist'][0] # Process each annotation for anno in anno_list: # Get image info img_name = anno.image.name[0][0] img_path = os.path.join(IMAGE_DIR, img_name) with Image.open(img_path) as img: img_w, img_h = img.size # Create label file label_file = os.path.join(LABEL_OUTPUT_DIR, f"{os.path.splitext(img_name)[0]}.txt") with open(label_file, 'w') as f: # Skip samples without joint annotations if not hasattr(anno, 'annopoints') or anno.annopoints.size == 0: continue # Extract joint data points = anno.annopoints[0][0].point for point in points: joint_idx = point.id[0][0] - 1 # MPI uses 1-indexed IDs, convert to 0-indexed x = point.x[0][0] y = point.y[0][0] # Skip unannotated joints if x == 0 or y == 0: continue # Normalize to YOLO format x_center = x / img_w y_center = y / img_h box_w = JOINT_BOX_PIXELS / img_w box_h = JOINT_BOX_PIXELS / img_h # Write to file f.write(f"{JOINT_CLASS_MAP[joint_idx]} {x_center:.6f} {y_center:.6f} {box_w:.6f} {box_h:.6f}\n")
- Adjust Joint Box Size: The fixed
JOINT_BOX_PIXELSvalue works for most cases, but you can make it adaptive (e.g., 1% of the image's shorter side) if you have varying image resolutions. - YOLO Keypoint Mode: If using YOLOv8/v9's keypoint detection, modify the script to output raw normalized coordinates instead of boxes. The format would be
class_id x1 y1 x2 y2 ... xn ynfor all joints in one sample. - Validate Labels: Spot-check a few label files against their corresponding images to ensure coordinates are correctly normalized and no missing annotations are included.
- MATLAB 7.3+ Files: If
scipy.io.loadmatfails, useh5pyinstead. Note that h5py uses slightly different indexing (e.g.,mat_data['imgpaths'][()]instead ofmat_data['imgpaths']).
内容的提问来源于stack exchange,提问作者Richard Price-Jones

