文件哈希计算脚本出现FileNotFoundError问题求助
问题描述
使用Python脚本计算目录中所有文件的哈希值时,部分随机文件会抛出FileNotFoundError: [Errno 2] No such file or directory错误。已使用文件夹绝对路径,脚本曾成功处理20个文件夹共约100GB的文件,说明脚本逻辑本身可用。出错的文件名无异常(例如title65name77.dwg),文件非空且可在对应程序中正常打开。
原脚本
import csv import hashlib import os folder_path = input('What folder do you want to get hashes from? ') output_file = input('What do you want the output CSV file to be named? (include .csv at the end of the file name)') # Generate hashes of files def hash_file(file_path): hasher = hashlib.sha256() with open(file_path, 'rb') as f: for chunk in iter(lambda: f.read(4096), b''): hasher.update(chunk) return hasher.hexdigest() # Record hashes & filepaths def record_hashes(folder_path, output_file): hashes = {} for root, _, files in os.walk(folder_path): for file in files: file_path = os.path.join(root, file) file_hash = hash_file(file_path) hashes[file_path] = file_hash # Write hashes and file paths to a CSV file with open(output_file, 'w', newline='') as f: writer = csv.writer(f) writer.writerow(['File Path', 'Hash']) for file_path, file_hash in hashes.items(): writer.writerow([file_path, file_hash]) record_hashes(folder_path, output_file) print("Done recording hashes of files in", folder_path)
解决方案
- 添加文件存在性校验:
os.walk遍历到文件后,到实际读取前,文件可能被删除、移动或重命名,导致路径失效。在处理前先检查文件是否存在且为普通文件,排除符号链接、临时文件等特殊情况。 - 单个文件异常捕获:对每个文件的哈希计算过程添加异常捕获,避免单个文件出错导致整个脚本终止,同时记录错误信息便于排查。
- 区分不同错误类型:将不存在、权限不足、读取失败等错误分类记录,方便后续定位问题。
修改后的脚本
import csv import hashlib import os folder_path = input('请输入要计算哈希值的文件夹路径:') output_file = input('请输入输出CSV文件名(需包含.csv后缀):') error_log = output_file.replace('.csv', '_errors.csv') # 自动生成错误日志文件名 def hash_file(file_path): hasher = hashlib.sha256() try: with open(file_path, 'rb') as f: for chunk in iter(lambda: f.read(4096), b''): hasher.update(chunk) return hasher.hexdigest(), None except Exception as e: return None, str(e) def record_hashes(folder_path, output_file, error_log): hashes = [] errors = [] for root, _, files in os.walk(folder_path): for file in files: file_path = os.path.join(root, file) # 先校验文件状态 if not os.path.exists(file_path): errors.append([file_path, '文件不存在']) continue if not os.path.isfile(file_path): errors.append([file_path, '非普通文件(如目录、符号链接)']) continue # 计算哈希并处理异常 file_hash, err_msg = hash_file(file_path) if file_hash: hashes.append([file_path, file_hash]) else: errors.append([file_path, err_msg]) # 写入正常结果到CSV with open(output_file, 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['文件路径', 'SHA256哈希值']) writer.writerows(hashes) # 写入错误日志(如果有错误) if errors: with open(error_log, 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['文件路径', '错误信息']) writer.writerows(errors) print(f"处理完成:成功计算 {len(hashes)} 个文件的哈希值,{len(errors)} 个文件出错,详情见 {error_log}") else: print(f"处理完成:所有 {len(hashes)} 个文件的哈希值已计算并写入 {output_file}") record_hashes(folder_path, output_file, error_log)
内容的提问来源于stack exchange,提问作者Tanner
相关产品推荐
相关产品推荐

