如何优化跨磁盘文件存在性检查脚本以提升运行速度
优化文件存在性检查脚本的速度方案
问题背景
我需要检查驱动器A中的文件是否存在于驱动器B中,写了如下Python脚本,但运行速度极慢,要检查242603个文件(总计145GB),求提速方案:
import os import subprocess # drive A SOURCE_PATH = '/media/username/8e223d5b-2755-4e9f-a2f6-fac5e762e836/username' # drive B DESTINY_PATH = '/home/username/' SUCCESS_CODE = 0 if __name__ == '__main__': source_file = '' destiny_file = '' for source_actualdir, source_subdir, source_dirFiles in os.walk(SOURCE_PATH): for source_filename in source_dirFiles: for destiny_actualdir, destiny_subdir, destiny_dirFiles in os.walk(DESTINY_PATH): for destiny_filename in destiny_dirFiles: source_file = os.path.join(source_actualdir, source_filename) destiny_file = os.path.join(destiny_actualdir, destiny_filename) response = subprocess.run(['diff', '-s', f'{source_file}', f'{destiny_file}'], capture_output=True) if response.returncode == SUCCESS_CODE: print(f'Coincidencia {source_file} == {destiny_file}') break print(f'File {source_file} is missing in {DESTINY_PATH}')
原脚本慢的核心原因
- 嵌套遍历灾难:对每个源文件都重新遍历整个目标驱动器,时间复杂度是O(M*N)(M为源文件数,N为目标文件数),24万级别的文件量直接导致指数级耗时。
- 频繁启动进程:每个文件对比都调用
diff子进程,进程启动的开销远大于文件内容对比本身。 - 逻辑漏洞:找到匹配后仅跳出内层文件名循环,外层目标目录遍历仍在继续;且不管是否匹配成功,最后都会打印“文件缺失”,逻辑错误。
提速方案
1. 先构建目标文件索引(核心优化)
一次性读取目标驱动器的所有文件信息并存储为字典,后续直接通过字典查询,将时间复杂度降到O(M):
import os import hashlib def build_target_index(target_path): # 用(文件名, 文件大小)作为键,避免同名不同文件误判;值为文件完整路径 target_index = {} for root, _, files in os.walk(target_path): for fname in files: full_path = os.path.join(root, fname) try: file_size = os.path.getsize(full_path) key = (fname, file_size) if key not in target_index: target_index[key] = full_path except OSError: # 跳过无法访问的文件 continue return target_index SOURCE_PATH = '/media/username/8e223d5b-2755-4e9f-a2f6-fac5e762e836/username' DESTINY_PATH = '/home/username/' if __name__ == '__main__': target_index = build_target_index(DESTINY_PATH) print("目标文件索引构建完成") for root, _, files in os.walk(SOURCE_PATH): for fname in files: source_full = os.path.join(root, fname) try: source_size = os.path.getsize(source_full) key = (fname, source_size) if key in target_index: target_full = target_index[key] # 用Python内置哈希对比,比调用diff快数倍 def calc_file_hash(file_path, block_size=65536): hasher = hashlib.blake2b() # blake2b比MD5更快且更安全 with open(file_path, 'rb') as f: while chunk := f.read(block_size): hasher.update(chunk) return hasher.hexdigest() source_hash = calc_file_hash(source_full) target_hash = calc_file_hash(target_full) if source_hash == target_hash: print(f'匹配成功:{source_full} == {target_full}') else: print(f'同名同大小但内容不同:{source_full} vs {target_full}') else: print(f'目标路径中缺失文件:{source_full}') except OSError: print(f'无法访问源文件:{source_full}')
2. 额外优化点
- 并行处理:如果磁盘IO允许,用
concurrent.futures开启多线程/多进程并行计算哈希,进一步提升速度(注意不要开太多进程,避免磁盘瓶颈)。 - 跳过无意义对比:先通过文件大小快速过滤,大小不同的直接跳过哈希计算。
- 缓存哈希值:如果目标目录存在多个同名同大小的文件,缓存它们的哈希值避免重复计算。
3. 直接用系统工具(无需自定义脚本)
如果不需要定制逻辑,用系统自带工具速度更快:
- rsync:执行
rsync -avn --delete SOURCE_PATH DESTINY_PATH,-n为模拟运行,会列出所有差异文件,速度极快。 - diff:执行
diff -rq SOURCE_PATH DESTINY_PATH,-r递归对比,-q仅输出差异,比手动嵌套调用diff高效得多。 - fdupes:专门用于查找重复文件的工具,可直接对比两个目录的重复项。
内容的提问来源于stack exchange,提问作者madtyn
相关产品推荐
相关产品推荐

