如何实现包含历史抽样结果的递进式随机文件抽取?
The issue with your original code is that each call to random.sample picks files independently—there’s no link between the first 25% selection and the second 50% selection, so overlap isn’t guaranteed. To fix this, we need to track the files selected in the first step and explicitly include them when making the second selection.
Here’s a revised approach that addresses this:
Step 1: Modify the Function to Reuse Selected Files
We’ll update the function to accept an optional list of files that must be included in the selection. This way, we can pass the 25% selected files when we need to build the 50% set.
import os import random import shutil def select_and_copy_files(data_path, out_path, percent, include_files=None): # Get all files in the source directory all_files = os.listdir(data_path) if include_files is not None: # Filter to only include files that actually exist in the source valid_included = [f for f in include_files if f in all_files] # Get files not in the included list remaining_files = [f for f in all_files if f not in valid_included] # Calculate how many more files we need to reach the desired percentage total_needed = round(len(all_files) * percent) additional_needed = max(total_needed - len(valid_included), 0) # Pick additional files from the remaining pool additional_files = random.sample(remaining_files, additional_needed) # Combine included and additional files selected_files = valid_included + additional_files else: # Normal selection: pick the desired percentage from all files total_needed = round(len(all_files) * percent) selected_files = random.sample(all_files, total_needed) # Create output folder if it doesn't exist os.makedirs(out_path, exist_ok=True) # Copy each selected file to the output folder for file_name in selected_files: src_file = os.path.join(data_path, file_name) dst_file = os.path.join(out_path, file_name) shutil.copy(src_file, dst_file) # Return the list of selected files for reuse return selected_files
Step 2: Execute the Selections
Now we can run the two steps in sequence, reusing the files from the first selection in the second:
# Step 1: Select 25% of files and copy to the '75' folder selected_25_percent = select_and_copy_files("Data", "75", 0.25) # Step 2: Select 50% of files (including the 25% from step 1) and copy to the '50' folder select_and_copy_files("Data", "50", 0.5, include_files=selected_25_percent)
How This Works:
- First Call: We pick 25% of all files randomly, copy them to the
75folder, and store the list of these files inselected_25_percent. - Second Call: We pass the
selected_25_percentlist to the function. The function first includes all those files, then randomly picks enough additional files from the remaining pool to reach 50% of the total files. This guarantees the 50% set includes the original 25%.
Edge Case Handling:
- If the included files already make up more than the desired percentage, the function will just use the included files (no additional files are picked).
- The
os.makedirs(out_path, exist_ok=True)ensures the output folder is created if it doesn’t exist, and doesn’t throw an error if it does.
内容的提问来源于stack exchange,提问作者Mejdi Dallel

