Python文件完整性监控:对比baseline哈希与目录文件哈希的方法
问题描述
我正在用Python开发一款文件完整性监控器,已实现遍历指定目录内所有文件、计算文件内容的SHA512哈希值,并将文件路径与对应哈希值存储到baseline.txt的功能。现在需要完善comp_baseline函数,实现以下功能:
- 将
baseline.txt中的数据解析为以文件路径为键、哈希值为值的字典 - 对比目录中文件的当前哈希值与字典中对应路径的哈希值,若哈希不匹配则发出告警
- 同时检测文件新增或删除的情况
当前代码存在逻辑错误,无法完成上述功能,baseline.txt的格式如下:
/home/kali/Documents/test_files/b.txt | c49b73859752f36533fe7efe8994a697f88b1f9fac06003ca2e7d5d1d97ddb230bebe54fc36610d105625862998126ff5974b40c322d719fd706c1db8d503958 /home/kali/Documents/test_files/e.txt | 2d32d7704de22f9016fc75269ee54a4576864e9983aeb187fe06bf00113e5a01b5c93b8123a1e20bcc1c213afe09ed70e7bf0e4b49df80a046b8f91a7daa0b17 /home/kali/Documents/test_files/a.txt | a07b2986313eff8ebe7d73d4f2432bd9c2a7d7c43867cfcb0712f30868302921c5b09d73cd823c6d15aaa6e66ecd024b11d1f09bdf95178031cfd539621b24b0 /home/kali/Documents/test_files/d.txt | febdef5ebb1eec24417b9341f42f3259f4e5918ebb6993fbd362574a8305b03b46da5046b14e7761ede08cb9c4176e89752d525c2917bf701e139397bc561040 /home/kali/Documents/test_files/c.txt | bc7f9d6af7a95e4ca293bc96e19f8047faa7ceefc516043fc16c29ee673855bda84c84dc57cb0f285301dd21f3fbada2a47548f2fbc03187d95c78eb4612822c
修正后的
comp_baseline函数 import os import hashlib import time # 根据实际情况修改全局变量 directory = "/home/kali/Documents/test_files" BUF_SIZE = 65536 # 64KB块大小,优化大文件读取效率 def comp_baseline(): base_dict = {} baseline_path = "baseline.txt" # 加载基线文件 if not os.path.exists(baseline_path): print("Baseline file does not exist in local directory.") return with open(baseline_path, 'r') as f: for line in f: line = line.strip() if not line: continue # 跳过空行 try: key, value = line.split(' | ') base_dict[key] = value.strip() except ValueError: print(f"Invalid line format in baseline: {line}") continue print("Loaded baseline successfully:") for path, hash_val in base_dict.items(): print(f"{path} | {hash_val[:10]}...") # 只打印哈希前10位,避免输出过长 # 持续监控文件状态 while True: time.sleep(1) current_files = set() # 遍历当前目录下的所有文件 for filename in os.listdir(directory): fn = os.path.join(directory, filename) # 跳过目录,只处理文件 if not os.path.isfile(fn): continue current_files.add(fn) # 计算当前文件的SHA512哈希 try: filehash = hashlib.sha512() # 每次计算新建哈希对象,避免残留旧数据 with open(fn, 'rb') as f2: while True: data = f2.read(BUF_SIZE) if not data: break filehash.update(data) current_hash = filehash.hexdigest() # 对比基线哈希值 if fn in base_dict: if current_hash != base_dict[fn]: print(f"⚠️ ALERT: File {fn} has been modified!") print(f" Old hash: {base_dict[fn][:10]}...") print(f" New hash: {current_hash[:10]}...") else: print(f"⚠️ ALERT: New file detected - {fn}") except Exception as e: print(f"Error processing file {fn}: {str(e)}") # 检测基线中存在但已被删除的文件 deleted_files = base_dict.keys() - current_files for deleted_fn in deleted_files: print(f"⚠️ ALERT: File {deleted_fn} has been deleted!")
关键修改说明
- 哈希对象重置:每次计算文件哈希时重新创建
hashlib.sha512()对象,避免之前计算的哈希值残留导致结果错误 - 修复存在性判断:用
fn in base_dict替代原代码中错误的fn in base_dict[key],正确判断文件是否在基线中 - 完整哈希对比:读取完整文件后计算哈希值,再与基线中的值进行对比,不匹配则触发修改告警
- 新增/删除检测:通过集合对比,检测出基线中没有的新文件、以及基线存在但已被删除的文件并告警
- 跳过目录:添加
os.path.isfile(fn)判断,避免遍历目录时出现读取错误 - 异常处理:捕获文件读取过程中的异常,防止单个文件处理失败导致整个监控中断
- 格式容错:处理基线文件中的空行和格式错误行,提高鲁棒性
内容的提问来源于stack exchange,提问作者AerialSong
相关产品推荐
相关产品推荐

