You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python os.walk()引用subdirs参数时运行结果异常问题求助

问题原因说明

核心错误

你第一版代码的异常是循环嵌套逻辑错误导致的,和文件系统、os.walk的实现无关:

  • os.walk每轮迭代返回的三个变量分别是:当前遍历目录路径root_path、当前目录下直属子目录列表subdirs、当前目录下直属文件列表files,它会自动递归遍历所有子目录,不需要你手动迭代subdirs触发递归
  • 你错误地把「遍历当前目录文件、筛选txt」的逻辑写在了「遍历subdirs」的循环内部,直接引发两个问题:
    • 当前目录下有N个直属子目录时,当前目录的文件会被重复扫描N次:比如测试目录中dir1下有2个直属子目录,所以dir1下的2个txt被重复输出了2次,和你第一版运行结果完全吻合
    • 没有直属子目录的目录,内部的文件会完全漏扫:比如dir1/subdir1、dir1/subdir2下没有子目录,for _ in subdirs循环根本不会执行,嵌套在内的文件扫描逻辑完全不会触发,因此这些目录下的txt全部没有出现在第一版的输出结果中

正确的进度条实现方案

如果要实现基于扫描目录数的进度条,你可以先做一次预遍历统计总目录数,再做第二次正式扫描:

import os

def count_total_dirs(start_from):
    total = 0
    for root_path, subdirs, files in os.walk(start_from):
        total += 1
    return total

def scan_for_txt_files(start_from):
    total_dirs = count_total_dirs(start_from)
    scanned_dirs = 0
    for root_path, subdirs, files in os.walk(start_from):
        scanned_dirs +=1
        # 此处更新进度条,进度=scanned_dirs / total_dirs
        for this_file in files:
            ext = str.lower(os.path.splitext(this_file)[1]).replace('.', '')
            if ext == 'txt':
                print(f'{os.path.join(root_path, this_file)}')

内容的提问来源于stack exchange,提问作者Alan B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 11:06:00