添加SHA-256哈希后,Python目录遍历脚本仅显示少量文件
问题分析与修复
你的脚本只显示少数文件的核心原因是缩进错误:哈希计算和表格行添加的代码没有包含在遍历单个文件的循环体内,导致每个目录下仅处理最后一个文件。
具体问题点
原代码中,with open(...)和tbl.add_row()位于for nextFile in fileList:循环的外部(循环结束后才执行),所以每遍历完一个目录下的所有文件,只会对最后一个文件计算哈希并添加到表格,最终表格里只有每个目录的最后一个文件数据。
修复后的完整代码
# Python Standard Libraries import os import hashlib import sys import time # Python 3rd Party Libraries from prettytable import PrettyTable # pip install prettytable # Local Functions def GetFileMetaData(fileName): # 获取文件系统元数据 try: metaData = os.stat(fileName) fileSize = metaData.st_size timeLastAccess = metaData.st_atime timeLastModified = metaData.st_mtime timeCreated = metaData.st_ctime macTimeList = [timeLastModified, timeCreated, timeLastAccess] return True, None, fileSize, macTimeList except Exception as err: return False, str(err), None, None # 脚本开始 tbl = PrettyTable(['FilePath','FileSize','UTC-Modified', 'UTC-Accessed', 'UTC-Created', 'SHA-256 HASH']) # 目标目录输入验证 while True: targetFolder = input("Enter Target Folder: ") if os.path.isdir(targetFolder): break else: print("\nInvalid Folder ... Please Try Again") print("Walking: ", targetFolder, "\n") for currentRoot, dirList, fileList in os.walk(targetFolder): for nextFile in fileList: fullPath = os.path.join(currentRoot, nextFile) absPath = os.path.abspath(fullPath) success, errInfo, fileSize, macList = GetFileMetaData(absPath) if success: # 转换为可读UTC时间 modTime = time.strftime("%Y-%m-%d %H:%M:%S", time.gmtime(macList[0])) accTime = time.strftime("%Y-%m-%d %H:%M:%S", time.gmtime(macList[1])) creTime = time.strftime("%Y-%m-%d %H:%M:%S", time.gmtime(macList[2])) # 哈希计算(移到文件循环内,且添加异常处理) hexDigest = "计算失败" try: # 大文件分块读取,避免内存占用过高 sha256Obj = hashlib.sha256() with open(absPath, 'rb') as target: while chunk := target.read(4096): sha256Obj.update(chunk) hexDigest = sha256Obj.hexdigest() except Exception as e: hexDigest = f"错误: {str(e)}" # 添加行到表格(移到文件循环内) tbl.add_row([absPath, fileSize, modTime, accTime, creTime, hexDigest]) tbl.align = "l" print(tbl.get_string(sortby="FileSize", reversesort=True)) print("\nScript-End\n")
关键修改说明
- 缩进修正:将哈希计算和
tbl.add_row()代码缩进,放入for nextFile in fileList:循环内部,确保每个文件都被处理。 - 大文件优化:将一次性读取文件内容改为分块读取(每次4096字节),避免处理大文件时内存溢出。
- 异常处理增强:给文件读取和哈希计算添加异常捕获,避免单个文件处理失败导致整个脚本终止,同时记录错误信息。
- 冗余代码移除:删除了多余的
fileSize = os.path.getsize(absPath),因为GetFileMetaData已经返回了正确的文件大小。
内容的提问来源于stack exchange,提问作者Melissa Zimmerman
相关产品推荐
相关产品推荐

