犬种识别CNN模型训练:如何将Stanford Dogs Dataset重构为Train/Test按犬种分类的文件夹结构
Alright, let's fix this folder structure for your ResNet training—Stanford Dogs Dataset actually comes with official train/test split files, so we can automate this instead of doing it manually. Here's exactly what to do in your Colab notebook:
Step 1: Confirm Your Raw Dataset Structure
First, make sure your extracted dataset has these key components in its root folder (let's assume it's stored at /content/stanford_dogs):
Images/: Contains subfolders named liken02085620-Chihuahua, each holding images for that specific breed.train_list.txt: Official list of training images (each line includes the relative path to an image and a label).test_list.txt: Official list of test images, formatted the same way as the training list.
Step 2: Run the Python Script to Restructure
Copy this code into a Colab cell, adjust the paths if your dataset is stored elsewhere, and execute it:
import os import shutil # -------------------------- # Update these paths if needed # -------------------------- base_raw_dir = "/content/stanford_dogs" # Path to your extracted dataset target_split_dir = "/content/stanford_dogs_split" # Where the structured dataset will be saved # -------------------------- # Set up source and destination paths # -------------------------- raw_images_dir = os.path.join(base_raw_dir, "Images") train_split_file = os.path.join(base_raw_dir, "train_list.txt") test_split_file = os.path.join(base_raw_dir, "test_list.txt") train_dest_dir = os.path.join(target_split_dir, "train") test_dest_dir = os.path.join(target_split_dir, "test") # Create base folders if they don't exist os.makedirs(train_dest_dir, exist_ok=True) os.makedirs(test_dest_dir, exist_ok=True) # -------------------------- # Function to split and copy images # -------------------------- def organize_images(split_file_path, destination_dir): with open(split_file_path, "r") as file: all_lines = file.readlines() for line in all_lines: # Split line to get the relative image path (ignore the label at the end) image_rel_path = line.strip().split()[0] breed_folder_name, image_filename = image_rel_path.split("/") # Build source and destination paths source_image_path = os.path.join(raw_images_dir, breed_folder_name, image_filename) breed_dest_folder = os.path.join(destination_dir, breed_folder_name) # Create breed folder in target if it doesn't exist os.makedirs(breed_dest_folder, exist_ok=True) dest_image_path = os.path.join(breed_dest_folder, image_filename) # Copy the image to the target location shutil.copy(source_image_path, dest_image_path) # -------------------------- # Run the organization for train and test sets # -------------------------- print("Organizing training set...") organize_images(train_split_file, train_dest_dir) print("Organizing test set...") organize_images(test_split_file, test_dest_dir) print("Done! Check the structured dataset at:", target_split_dir)
Step 3: Verify the Result
After running the script, navigate to /content/stanford_dogs_split in Colab's file explorer. You'll see:
train/: Contains subfolders for each dog breed, with all corresponding training images inside.test/: Mirror structure, but populated with test images.
This structure is perfect for using with Keras' ImageDataGenerator or PyTorch's ImageFolder—both tools will automatically map each breed folder to a unique class label, which is exactly what you need for your ResNet training pipeline.
A quick heads-up: The dataset has ~20k images total, so the copy process might take a minute or two depending on Colab's current speed. No need to panic if it doesn't finish instantly!
内容的提问来源于stack exchange,提问作者Ayushya Pare

