You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按比例/数量随机复制匹配的图片与文本文件至多目录?

Nice start with the batch copy functionality! Let's expand it to meet your two sampling requirements—handling paired .png and .txt files, ensuring randomness, and avoiding duplicate files across samples. Here's a complete, modular solution:

1. Core Setup & Helper Function

First, let's create a reusable helper function to copy paired files (both the .png and its matching .txt) and handle directory creation automatically:

import glob
import os
import shutil
import random

def copy_paired_files(basename, source_dir, dest_dir):
    """Copy both .png and .txt files with the given basename to the destination directory"""
    # Create destination directory if it doesn't exist
    os.makedirs(dest_dir, exist_ok=True)
    
    # Copy the PNG file
    png_path = os.path.join(source_dir, f"{basename}.png")
    if os.path.isfile(png_path):
        shutil.copy2(png_path, dest_dir)
    
    # Copy the corresponding TXT file
    txt_path = os.path.join(source_dir, f"{basename}.txt")
    if os.path.isfile(txt_path):
        shutil.copy2(txt_path, dest_dir)

2. Feature 1: Random Percentage Sampling

This function will randomly select a specified percentage of paired files, with unique results on every run (thanks to Python's default time-based random seed):

def sample_by_percentage(source_dir, dest_dir, percentage):
    # Get all unique basenames from PNG files (each maps to a matching TXT)
    png_files = glob.glob(os.path.join(source_dir, "*.png"))
    basenames = [os.path.splitext(os.path.basename(file))[0] for file in png_files]
    
    total_files = len(basenames)
    sample_count = int(total_files * percentage / 100)
    
    # Handle edge case where percentage is too small
    if sample_count == 0:
        print(f"Warning: {percentage}% of {total_files} files results in 0 samples")
        return
    
    # Randomly select unique basenames (no duplicates in the sample)
    sampled_basenames = random.sample(basenames, sample_count)
    
    # Copy all paired files for the sampled basenames
    for basename in sampled_basenames:
        copy_paired_files(basename, source_dir, dest_dir)
    
    print(f"Successfully copied {sample_count} paired files to {dest_dir}")

# Example usage: Extract 40% of files to 'percentage_sample/'
sample_by_percentage('all_the_content/', 'percentage_sample/', 40)

3. Feature 2: Split into Multiple Unique Samples

This function splits all paired files into a specified number of separate samples, ensuring no file appears in more than one sample:

def split_into_samples(source_dir, sample_count, base_dest_dir="samples/"):
    # Get all unique basenames
    png_files = glob.glob(os.path.join(source_dir, "*.png"))
    basenames = [os.path.splitext(os.path.basename(file))[0] for file in png_files]
    total_files = len(basenames)
    
    # Validate input
    if sample_count <= 0:
        print("Error: Sample count must be a positive number")
        return
    if total_files < sample_count:
        print(f"Warning: More samples ({sample_count}) than files ({total_files})—some samples will be empty")
    
    # Shuffle basenames to ensure random distribution
    random.shuffle(basenames)
    
    # Split into roughly equal chunks (handle uneven division)
    chunk_size = total_files // sample_count
    remainder = total_files % sample_count
    
    samples = []
    start_idx = 0
    for i in range(sample_count):
        # Add one extra file to the first 'remainder' samples for even distribution
        end_idx = start_idx + chunk_size + (1 if i < remainder else 0)
        samples.append(basenames[start_idx:end_idx])
        start_idx = end_idx
    
    # Copy each sample to its own directory
    for sample_idx, sample_basenames in enumerate(samples, 1):
        sample_dir = os.path.join(base_dest_dir, f"sample_{sample_idx}")
        for basename in sample_basenames:
            copy_paired_files(basename, source_dir, sample_dir)
        print(f"Sample {sample_idx}: Copied {len(sample_basenames)} paired files to {sample_dir}")

# Example usage: Split into 3 samples stored in 'samples/'
split_into_samples('all_the_content/', 3)

Key Notes

  • Randomness: By default, every run will produce different results (Python uses system time to seed the randomizer). If you need consistent results for testing, add random.seed(42) (or any fixed number) at the top of your script.
  • Missing Files: The code will skip copying .txt files if they don't exist for a given .png, no errors will be thrown.

内容的提问来源于stack exchange,提问作者Jess

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:48:25