如何基于Python文件夹名称列表从指定目录中筛选目标文件夹并提取其中文件
Fixing Directory Traversal & File Labeling Issues
Hey, let's work through fixing your code step by step. I spot several key issues that are making it misbehave, plus some optimizations to make it run smoother and do exactly what you need.
First, let's break down why your original code isn't working as expected:
- Path Escaping Problems: In Windows, backslashes
\act as escape characters. So something like\Trainmight get parsed incorrectly. You need raw strings or double backslashes to avoid this. - Redundant Traversal & Duplicate Entries: You’re looping through every item in
df_train_pos_listand running a fullos.walk()each time. That means every file gets added totrain_imagesmultiple times (once per item in your list)—total overkill. - Broken Path Concatenation: Just tacking the filename onto the root path skips subdirectory structure, leading to invalid paths like
D:\Arm C Deep Learning\SH_OCTAPUS\Trainfile.jpg. - Label Logic Errors: Files in non-target folders get labeled 0 once for every item in
df_train_pos_list, creating massive duplicates. Target folder files also get re-added repeatedly.
Here's the corrected, optimized version of your code:
import os import numpy as np train_images = [] train_labels = [] # Use a raw string for Windows paths to avoid escape character mishaps root_directory = r'D:\Arm C Deep Learning\SH_OCTAPUS\Train' # Convert your folder list to a set for near-instant membership checks target_folders = set(df_train_pos_list) # Only traverse the directory once to avoid duplicate processing for current_path, subfolders, files in os.walk(root_directory): # Grab the name of the current folder (the last segment of the path) current_folder = os.path.basename(current_path) # Check if this folder is in our target list is_target = current_folder in target_folders for file in files: # Build a valid full path using os.path.join (handles separators automatically) full_file_path = os.path.join(current_path, file) train_images.append(full_file_path) # Assign label 1 for target folders, 0 otherwise train_labels.append(1 if is_target else 0) # Convert labels to a numpy array train_labels = np.asarray(train_labels, dtype=np.int64) print(train_labels) # Fix the return_counts parameter (needs a boolean, not a string) unique_labels, label_counts = np.unique(train_labels, return_counts=True) print(f"Label distribution: {dict(zip(unique_labels, label_counts))}")
Key Improvements Explained:
- Safe Path Handling: Raw strings (
r'path') ensure Windows paths are parsed correctly without escape character confusion. - Single Directory Traversal: We run
os.walk()once, so every file is processed exactly once—no more duplicate entries cluttering your lists. - Faster Folder Checks: Sets make
inchecks nearly instant (O(1) vs. O(n) for lists), which is a big win if your folder list is large. - Valid Paths:
os.path.join()automatically handles path separators, so you’ll always get a correct full path, even for files nested in subdirectories. - Clean Label Logic: We check once per folder if it’s a target, then apply the label to all its files—no repeated checks or redundant entries.
- Fixed
np.uniqueCall: Thereturn_countsparameter expects a boolean (True), not a string ('TRUE'), which was causing an error in your original code.
内容的提问来源于stack exchange,提问作者sum87
相关产品推荐
相关产品推荐

