如何优化Python中十万级文件的复制效率?
解决大量文件批量复制的高效方法
针对10万级文件的复制需求,直接循环调用shutil.copy()确实会因频繁系统调用导致效率低下,以下是几个更高效的替代方案:
调用系统原生
cp命令(推荐)
Unix系统的cp -r是经过高度优化的原生工具,处理大量文件的速度和终端执行一致,远快于Python循环调用。可以用subprocess模块直接执行:import subprocess # 递归复制整个目录 subprocess.run(['cp', '-r', '/path/to/source', '/path/to/destination'], check=True)该方案完全复用系统级优化,适合无需额外Python逻辑的场景。
使用
shutil.copytree()
Python标准库的shutil.copytree()是专门用于递归复制目录的方法,底层实现比手动循环shutil.copy()更高效,还能自动处理目录结构:import shutil # Python 3.8+支持dirs_exist_ok参数,允许目标目录已存在 shutil.copytree('/path/to/source', '/path/to/destination', dirs_exist_ok=True)如果需要过滤特定文件,可通过
ignore参数实现,比如忽略所有.log文件:shutil.copytree(src, dest, ignore=shutil.ignore_patterns('*.log'), dirs_exist_ok=True)多进程批量复制(适合需自定义逻辑的场景)
如果必须逐个处理文件但想提升效率,可尝试多进程并行复制(注意:IO密集型任务提升幅度有限,需根据实际环境测试):import os import shutil from concurrent.futures import ProcessPoolExecutor def copy_file(file_pair): src_file, dest_file = file_pair os.makedirs(os.path.dirname(dest_file), exist_ok=True) shutil.copy2(src_file, dest_file) src_dir = '/path/to/source' dest_dir = '/path/to/destination' # 生成所有待复制的文件路径对 file_pairs = [] for root, _, files in os.walk(src_dir): for file in files: src_path = os.path.join(root, file) dest_path = os.path.join(dest_dir, os.path.relpath(src_path, src_dir)) file_pairs.append((src_path, dest_path)) # 并行复制,进程数可根据CPU核心数调整 with ProcessPoolExecutor(max_workers=4) as executor: executor.map(copy_file, file_pairs)
内容的提问来源于stack exchange,提问作者Dmitry Sokolov
相关产品推荐
相关产品推荐

