如何用Python递归或os.walk统计子目录文件年份并生成DataFrame
基于os.walk的子目录文件年份统计方案
目录结构
. ├── a.txt ├── b.txt ├── foo │ └── w.txt │ └── a.txt └── moo └── cool.csv └── bad.csv └── more └── wow.csv
需求说明
统计每个子目录下文件的年份数量,生成如下格式的Pandas DataFrame:
Subdir 2020 2021 2022 foo 0 1 1 moo 0 2 0 more 1 0 0
原递归实现的函数运行时导致内核崩溃,需改用os.walk实现。
原递归代码
import os import pandas as pd dir_path = 'S:\\Test' def getFiles(dir_path): contents = os.listdir(dir_path) # check if content is directory or not for file in contents: if os.path.isdir(os.path.join(dir_path, file)): # get everything inside subdirectory getFiles(dir_path = os.path.join(dir_path, file)) # it's a file else: # do something to get the year of the file and put it in a list or something # at the end create pandas data frame and return
解决方案代码
import os import pandas as pd from datetime import datetime dir_path = 'S:\\Test' # 定义需要统计的年份范围 target_years = [2020, 2021, 2022] # 初始化统计字典,键为子目录名,值为对应年份的计数 stats = {} # 使用os.walk遍历目录 for root, dirs, files in os.walk(dir_path): # 跳过根目录,只处理子目录 if root == dir_path: continue # 获取当前子目录的名称(取最后一级目录名) subdir_name = os.path.basename(root) # 初始化当前子目录的年份计数为0 year_counts = {year: 0 for year in target_years} for file in files: file_path = os.path.join(root, file) # 获取文件的修改时间(也可根据需求改为创建时间os.path.getctime) timestamp = os.path.getmtime(file_path) file_year = datetime.fromtimestamp(timestamp).year # 如果年份在目标范围内,计数加1 if file_year in year_counts: year_counts[file_year] += 1 # 将当前子目录的统计结果存入字典 stats[subdir_name] = year_counts # 转换为Pandas DataFrame并调整格式 df = pd.DataFrame.from_dict(stats, orient='index').reset_index() df.columns = ['Subdir'] + target_years # 按示例格式排序(可选,根据需求调整) df = df.set_index('Subdir').reindex(['foo', 'moo', 'more']).reset_index() print(df)
代码说明
- 用
os.walk替代递归,避免目录过深引发的内核崩溃问题,os.walk内部已优化遍历逻辑 - 通过
os.path.getmtime获取文件修改时间,如需统计创建时间可替换为os.path.getctime - 提前初始化目标年份的计数字典,确保每个年份都有统计项
- 最后将统计结果转换为指定格式的DataFrame,可根据实际需求调整子目录排序顺序
内容的提问来源于stack exchange,提问作者Eisen
相关产品推荐
相关产品推荐

