You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python递归或os.walk统计子目录文件年份并生成DataFrame

基于os.walk的子目录文件年份统计方案

目录结构

.
├── a.txt
├── b.txt
├── foo
│   └── w.txt
│   └── a.txt
└── moo
    └── cool.csv
    └── bad.csv
    └── more
        └── wow.csv

需求说明

统计每个子目录下文件的年份数量,生成如下格式的Pandas DataFrame:

Subdir 2020 2021 2022
foo    0    1    1
moo    0    2    0
more   1    0    0

原递归实现的函数运行时导致内核崩溃,需改用os.walk实现。

原递归代码

import os
import pandas as pd

dir_path = 'S:\\Test'

def getFiles(dir_path):
     contents = os.listdir(dir_path)
     # check if content is directory or not
     for file in contents:
          if os.path.isdir(os.path.join(dir_path, file)):
               # get everything inside subdirectory
               getFiles(dir_path = os.path.join(dir_path, file))
          # it's a file
          else:
               # do something to get the year of the file and put it in a list or something
     # at the end create pandas data frame and return

解决方案代码

import os
import pandas as pd
from datetime import datetime

dir_path = 'S:\\Test'
# 定义需要统计的年份范围
target_years = [2020, 2021, 2022]

# 初始化统计字典,键为子目录名,值为对应年份的计数
stats = {}

# 使用os.walk遍历目录
for root, dirs, files in os.walk(dir_path):
    # 跳过根目录,只处理子目录
    if root == dir_path:
        continue
    # 获取当前子目录的名称(取最后一级目录名)
    subdir_name = os.path.basename(root)
    # 初始化当前子目录的年份计数为0
    year_counts = {year: 0 for year in target_years}
    
    for file in files:
        file_path = os.path.join(root, file)
        # 获取文件的修改时间(也可根据需求改为创建时间os.path.getctime)
        timestamp = os.path.getmtime(file_path)
        file_year = datetime.fromtimestamp(timestamp).year
        # 如果年份在目标范围内,计数加1
        if file_year in year_counts:
            year_counts[file_year] += 1
    
    # 将当前子目录的统计结果存入字典
    stats[subdir_name] = year_counts

# 转换为Pandas DataFrame并调整格式
df = pd.DataFrame.from_dict(stats, orient='index').reset_index()
df.columns = ['Subdir'] + target_years
# 按示例格式排序(可选,根据需求调整)
df = df.set_index('Subdir').reindex(['foo', 'moo', 'more']).reset_index()

print(df)

代码说明

  • 用os.walk替代递归,避免目录过深引发的内核崩溃问题,os.walk内部已优化遍历逻辑
  • 通过os.path.getmtime获取文件修改时间,如需统计创建时间可替换为os.path.getctime
  • 提前初始化目标年份的计数字典,确保每个年份都有统计项
  • 最后将统计结果转换为指定格式的DataFrame,可根据实际需求调整子目录排序顺序

内容的提问来源于stack exchange,提问作者Eisen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 20:50:29