You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python自动创建序号递增文件夹并批量分类PDF文件

问题解决:按批次拆分PDF文件并自定义文件夹命名

需求说明

  • 将约3800个PDF文件按每50个一批拆分到文件夹中
  • 两种文件夹命名/指定方案:
    • 方案①:按数字序号(从1开始)创建并命名新文件夹
    • 方案②:将文件放入已提前创建好的空模板文件夹
  • 需要记录每个文件对应的所属文件夹,用于后续打印软件输出及文件列表生成

当前代码存在的问题

  • 现有拆分代码用批次内首个文件的名称命名文件夹,不符合序号命名需求
  • 尝试的批量重命名代码因缩进错误(os.rename语句在循环外部),仅会重命名最后一个文件夹,未生效

方案1:按数字序号创建并命名新文件夹

直接修改原拆分代码,用数字序号作为文件夹名称,同时自动生成文件归属记录:

import os
import shutil

source_directory = "//All-the-data"
destination_base_folder = '//All-the-data/sorted'
batch_size = 50
# 生成文件归属记录的路径
record_file_path = os.path.join(destination_base_folder, "file_batch_record.txt")

# 检查源目录是否存在
if os.path.exists(source_directory) and os.path.isdir(source_directory):
    # 仅获取源目录中的文件(过滤掉子文件夹),并按文件名排序
    files = [f for f in os.listdir(source_directory) if os.path.isfile(os.path.join(source_directory, f))]
    files.sort()

    batch_counter = 1
    # 清空并初始化记录文件
    with open(record_file_path, 'w', encoding='utf-8') as record_file:
        # 按批次遍历文件
        for i in range(0, len(files), batch_size):
            # 用带补零的数字序号命名文件夹,保证排序美观(如Batch_001、Batch_002)
            batch_folder_name = f"Batch_{batch_counter:03d}"
            batch_folder_path = os.path.join(destination_base_folder, batch_folder_name)
            # 创建文件夹(已存在则跳过)
            os.makedirs(batch_folder_path, exist_ok=True)

            # 复制当前批次的文件
            current_batch_files = files[i:i+batch_size]
            for file_name in current_batch_files:
                source_file = os.path.join(source_directory, file_name)
                dest_file = os.path.join(batch_folder_path, file_name)
                shutil.copy2(source_file, dest_file)
                # 记录文件与对应文件夹的关系
                record_file.write(f"{file_name}\t{batch_folder_name}\n")

            print(f"批次 {batch_counter} 已成功复制到 {batch_folder_path}")
            batch_counter += 1
    print(f"文件归属记录已保存至:{record_file_path}")
else:
    print("错误:源文件夹不存在 -> " + source_directory)

代码说明

  • 自动过滤源目录中的子文件夹,仅处理PDF文件
  • 文件夹命名采用Batch_001格式,补零确保文件夹按数字顺序排列
  • 生成file_batch_record.txt文本文件,每行记录文件名\t所属文件夹,方便后续查看文件归属

方案2:将文件放入已有的空模板文件夹

适用于已提前在目标目录下创建好空模板文件夹的场景(需保证模板数量足够:3800/50=76个):

import os
import shutil

source_directory = "//All-the-data"
destination_base_folder = '//All-the-data/sorted'
batch_size = 50
record_file_path = os.path.join(destination_base_folder, "file_batch_record.txt")

# 检查源目录是否存在
if os.path.exists(source_directory) and os.path.isdir(source_directory):
    files = [f for f in os.listdir(source_directory) if os.path.isfile(os.path.join(source_directory, f))]
    files.sort()

    # 获取目标目录下的所有空模板文件夹,并按名称排序
    template_folders = [f for f in os.listdir(destination_base_folder) if os.path.isdir(os.path.join(destination_base_folder, f))]
    template_folders.sort()

    # 计算所需批次数量,检查模板文件夹是否足够
    required_batch_count = (len(files) + batch_size - 1) // batch_size
    if len(template_folders) < required_batch_count:
        print(f"错误:模板文件夹数量不足,需要至少 {required_batch_count} 个,当前仅 {len(template_folders)} 个")
        exit()

    # 开始复制文件并记录归属
    with open(record_file_path, 'w', encoding='utf-8') as record_file:
        for batch_idx in range(required_batch_count):
            current_template_folder = template_folders[batch_idx]
            folder_path = os.path.join(destination_base_folder, current_template_folder)
            # 确定当前批次的文件范围
            start_idx = batch_idx * batch_size
            end_idx = start_idx + batch_size
            current_batch_files = files[start_idx:end_idx]

            for file_name in current_batch_files:
                source_file = os.path.join(source_directory, file_name)
                dest_file = os.path.join(folder_path, file_name)
                shutil.copy2(source_file, dest_file)
                record_file.write(f"{file_name}\t{current_template_folder}\n")

            print(f"批次 {batch_idx+1} 已成功复制到 {folder_path}")
    print(f"文件归属记录已保存至:{record_file_path}")
else:
    print("错误:源文件夹不存在 -> " + source_directory)

代码说明

  • 先获取并排序已有的模板文件夹,确保按顺序分配文件
  • 提前检查模板数量是否足够,避免中途出错
  • 同样生成文件归属记录,方便后续管理

修正你之前的批量重命名代码

你写的重命名代码因缩进错误导致失效,修正后版本:

def number_folders_chronologically(destination_folder):
    # 获取目标目录下的所有文件夹
    batch_folders = [folder for folder in os.listdir(destination_folder) if os.path.isdir(os.path.join(destination_folder, folder))]
    # 按文件夹创建时间排序
    batch_folders.sort(key=lambda x: os.path.getctime(os.path.join(destination_folder, x)))

    # 遍历并重命名每个文件夹
    for index, folder in enumerate(batch_folders, start=1):
        old_path = os.path.join(destination_folder, folder)
        new_folder_name = f"{index:03d}_{folder}"
        new_path = os.path.join(destination_folder, new_folder_name)
        # 注意:将os.rename放入循环内部,否则只会处理最后一个文件夹
        os.rename(old_path, new_path)

内容的提问来源于stack exchange,提问作者nirby

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 07:51:03