You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Python文件夹名称列表从指定目录中筛选目标文件夹并提取其中文件

Fixing Directory Traversal & File Labeling Issues

Hey, let's work through fixing your code step by step. I spot several key issues that are making it misbehave, plus some optimizations to make it run smoother and do exactly what you need.

First, let's break down why your original code isn't working as expected:

  • Path Escaping Problems: In Windows, backslashes \ act as escape characters. So something like \Train might get parsed incorrectly. You need raw strings or double backslashes to avoid this.
  • Redundant Traversal & Duplicate Entries: You’re looping through every item in df_train_pos_list and running a full os.walk() each time. That means every file gets added to train_images multiple times (once per item in your list)—total overkill.
  • Broken Path Concatenation: Just tacking the filename onto the root path skips subdirectory structure, leading to invalid paths like D:\Arm C Deep Learning\SH_OCTAPUS\Trainfile.jpg.
  • Label Logic Errors: Files in non-target folders get labeled 0 once for every item in df_train_pos_list, creating massive duplicates. Target folder files also get re-added repeatedly.

Here's the corrected, optimized version of your code:

import os
import numpy as np

train_images = []
train_labels = []
# Use a raw string for Windows paths to avoid escape character mishaps
root_directory = r'D:\Arm C Deep Learning\SH_OCTAPUS\Train'
# Convert your folder list to a set for near-instant membership checks
target_folders = set(df_train_pos_list)

# Only traverse the directory once to avoid duplicate processing
for current_path, subfolders, files in os.walk(root_directory):
    # Grab the name of the current folder (the last segment of the path)
    current_folder = os.path.basename(current_path)
    # Check if this folder is in our target list
    is_target = current_folder in target_folders
    
    for file in files:
        # Build a valid full path using os.path.join (handles separators automatically)
        full_file_path = os.path.join(current_path, file)
        train_images.append(full_file_path)
        # Assign label 1 for target folders, 0 otherwise
        train_labels.append(1 if is_target else 0)

# Convert labels to a numpy array
train_labels = np.asarray(train_labels, dtype=np.int64)
print(train_labels)
# Fix the return_counts parameter (needs a boolean, not a string)
unique_labels, label_counts = np.unique(train_labels, return_counts=True)
print(f"Label distribution: {dict(zip(unique_labels, label_counts))}")

Key Improvements Explained:

  • Safe Path Handling: Raw strings (r'path') ensure Windows paths are parsed correctly without escape character confusion.
  • Single Directory Traversal: We run os.walk() once, so every file is processed exactly once—no more duplicate entries cluttering your lists.
  • Faster Folder Checks: Sets make in checks nearly instant (O(1) vs. O(n) for lists), which is a big win if your folder list is large.
  • Valid Paths: os.path.join() automatically handles path separators, so you’ll always get a correct full path, even for files nested in subdirectories.
  • Clean Label Logic: We check once per folder if it’s a target, then apply the label to all its files—no repeated checks or redundant entries.
  • Fixed np.unique Call: The return_counts parameter expects a boolean (True), not a string ('TRUE'), which was causing an error in your original code.

内容的提问来源于stack exchange,提问作者sum87

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 11:53:00