You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Python中十万级文件的复制效率?

解决大量文件批量复制的高效方法

针对10万级文件的复制需求,直接循环调用shutil.copy()确实会因频繁系统调用导致效率低下,以下是几个更高效的替代方案:

  • 调用系统原生cp命令(推荐)
    Unix系统的cp -r是经过高度优化的原生工具,处理大量文件的速度和终端执行一致,远快于Python循环调用。可以用subprocess模块直接执行:

    import subprocess
    # 递归复制整个目录
    subprocess.run(['cp', '-r', '/path/to/source', '/path/to/destination'], check=True)
    

    该方案完全复用系统级优化,适合无需额外Python逻辑的场景。

  • 使用shutil.copytree()
    Python标准库的shutil.copytree()是专门用于递归复制目录的方法,底层实现比手动循环shutil.copy()更高效,还能自动处理目录结构:

    import shutil
    # Python 3.8+支持dirs_exist_ok参数,允许目标目录已存在
    shutil.copytree('/path/to/source', '/path/to/destination', dirs_exist_ok=True)
    

    如果需要过滤特定文件,可通过ignore参数实现,比如忽略所有.log文件:

    shutil.copytree(src, dest, ignore=shutil.ignore_patterns('*.log'), dirs_exist_ok=True)
    
  • 多进程批量复制(适合需自定义逻辑的场景)
    如果必须逐个处理文件但想提升效率,可尝试多进程并行复制(注意:IO密集型任务提升幅度有限,需根据实际环境测试):

    import os
    import shutil
    from concurrent.futures import ProcessPoolExecutor
    
    def copy_file(file_pair):
        src_file, dest_file = file_pair
        os.makedirs(os.path.dirname(dest_file), exist_ok=True)
        shutil.copy2(src_file, dest_file)
    
    src_dir = '/path/to/source'
    dest_dir = '/path/to/destination'
    
    # 生成所有待复制的文件路径对
    file_pairs = []
    for root, _, files in os.walk(src_dir):
        for file in files:
            src_path = os.path.join(root, file)
            dest_path = os.path.join(dest_dir, os.path.relpath(src_path, src_dir))
            file_pairs.append((src_path, dest_path))
    
    # 并行复制,进程数可根据CPU核心数调整
    with ProcessPoolExecutor(max_workers=4) as executor:
        executor.map(copy_file, file_pairs)
    

内容的提问来源于stack exchange,提问作者Dmitry Sokolov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 06:07:19