You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python删除待解密哈希文件中已存在于解密结果文件的条目

问题排查

核心错误点如下:

  • 错误使用w+模式打开待解密文件:w+模式会在打开文件时直接清空原有内容,还没读取原待解密哈希就把文件清了,这是文件变为0字节的直接原因。
  • 未读取待解密文件的原有内容:后续遍历逻辑错误遍历了已经读完的已解密文件对象,而没有读取待解密文件的内容,没有任何内容写入新文件,最终文件为空。
  • 匹配逻辑完全错误:判断条件word == hashes_to_delete_list是将单个哈希值和整个待删除列表做相等判断,条件永远不成立,就算有内容也不会触发写入逻辑。
  • 性能问题:用列表存储待删除哈希,嵌套循环匹配的时间复杂度是O(n*m),数十万条数据下运行效率极低,应该用集合存储实现O(1)时间复杂度的匹配。
  • 多余的文件关闭操作:with上下文管理器会自动处理文件关闭,后续手动调用close()属于冗余操作,极端情况下会触发异常。

修复方案

采用临时文件中转的方式处理,避免中途出错导致原文件损坏,完整可运行代码如下:

import os

file_with_hashes_found = "hashcat_test.potfile"
file_with_hashes_to_find = "hashes_test.txt"
temp_file = "hashes_test_temp.txt"
deleted_count = 0

# 读取所有已解密的哈希存入集合,去重+快速查找
decrypted_hashes = set()
print(f"DEV MESSAGE: 正在读取已解密文件 {file_with_hashes_found}")
with open(file_with_hashes_found, encoding="latin1") as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        hash_val = line.split(":", 1)[0]
        decrypted_hashes.add(hash_val)
print(f"DEV MESSAGE: 共加载 {len(decrypted_hashes)} 条已解密哈希")

# 过滤待解密文件,保留未解密的条目
print(f"DEV MESSAGE: 正在过滤待解密文件 {file_with_hashes_to_find}")
with open(file_with_hashes_to_find, "r", encoding="latin1") as f_in, open(temp_file, "w", encoding="latin1") as f_out:
    for line in f_in:
        line_stripped = line.strip()
        if not line_stripped:
            continue
        if line_stripped in decrypted_hashes:
            deleted_count += 1
            continue
        f_out.write(line)

# 用处理完成的临时文件替换原待解密文件
os.replace(temp_file, file_with_hashes_to_find)
print(f'DEV MESSAGE: 处理完成,共删除 {deleted_count} 条已解密哈希')

内容的提问来源于stack exchange,提问作者Mariusz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 06:24:00