You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

添加SHA-256哈希后,Python目录遍历脚本仅显示少量文件

问题分析与修复

你的脚本只显示少数文件的核心原因是缩进错误:哈希计算和表格行添加的代码没有包含在遍历单个文件的循环体内,导致每个目录下仅处理最后一个文件。

具体问题点

原代码中,with open(...)和tbl.add_row()位于for nextFile in fileList:循环的外部(循环结束后才执行),所以每遍历完一个目录下的所有文件,只会对最后一个文件计算哈希并添加到表格,最终表格里只有每个目录的最后一个文件数据。

修复后的完整代码

# Python Standard Libraries 
import os
import hashlib
import sys
import time

# Python 3rd Party Libraries
from prettytable import PrettyTable     # pip install prettytable

# Local Functions
def GetFileMetaData(fileName):
    # 获取文件系统元数据
    try:
        metaData         = os.stat(fileName)
        fileSize         = metaData.st_size
        timeLastAccess   = metaData.st_atime
        timeLastModified = metaData.st_mtime
        timeCreated      = metaData.st_ctime      
    
        macTimeList = [timeLastModified, timeCreated, timeLastAccess]
        return True, None, fileSize, macTimeList

    except Exception as err:
        return False, str(err), None, None

# 脚本开始
tbl = PrettyTable(['FilePath','FileSize','UTC-Modified', 'UTC-Accessed', 'UTC-Created', 'SHA-256 HASH'])

# 目标目录输入验证
while True:
    targetFolder = input("Enter Target Folder: ")
    if os.path.isdir(targetFolder):
        break
    else:
        print("\nInvalid Folder ... Please Try Again") 

print("Walking: ", targetFolder, "\n")

for currentRoot, dirList, fileList in os.walk(targetFolder):
    for nextFile in fileList:
        fullPath = os.path.join(currentRoot, nextFile)
        absPath  = os.path.abspath(fullPath)
        success, errInfo, fileSize, macList = GetFileMetaData(absPath)     
    
        if success:
            # 转换为可读UTC时间
            modTime = time.strftime("%Y-%m-%d %H:%M:%S", time.gmtime(macList[0]))
            accTime = time.strftime("%Y-%m-%d %H:%M:%S", time.gmtime(macList[1]))               
            creTime = time.strftime("%Y-%m-%d %H:%M:%S", time.gmtime(macList[2]))  

            # 哈希计算(移到文件循环内,且添加异常处理)
            hexDigest = "计算失败"
            try:
                # 大文件分块读取,避免内存占用过高
                sha256Obj = hashlib.sha256()
                with open(absPath, 'rb') as target:
                    while chunk := target.read(4096):
                        sha256Obj.update(chunk)
                hexDigest = sha256Obj.hexdigest()
            except Exception as e:
                hexDigest = f"错误: {str(e)}"
            
            # 添加行到表格(移到文件循环内)
            tbl.add_row([absPath, fileSize, modTime, accTime, creTime, hexDigest])

tbl.align = "l"
print(tbl.get_string(sortby="FileSize", reversesort=True))
print("\nScript-End\n")

关键修改说明

  • 缩进修正:将哈希计算和tbl.add_row()代码缩进,放入for nextFile in fileList:循环内部,确保每个文件都被处理。
  • 大文件优化:将一次性读取文件内容改为分块读取(每次4096字节),避免处理大文件时内存溢出。
  • 异常处理增强:给文件读取和哈希计算添加异常捕获,避免单个文件处理失败导致整个脚本终止,同时记录错误信息。
  • 冗余代码移除:删除了多余的fileSize = os.path.getsize(absPath),因为GetFileMetaData已经返回了正确的文件大小。

内容的提问来源于stack exchange,提问作者Melissa Zimmerman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 12:50:22