Python程序遍历rockyouu.txt时无法处理第230816行的问题求助
解决rockyouu.txt读取时的UnicodeDecodeError问题
问题根源
rockyou.txt原始数据集并非UTF-8编码,而是使用**Latin-1(ISO-8859-1)**或Windows-1252这类单字节编码。即使尝试转成UTF-8,某些无法被UTF-8兼容的字节依然会导致解码失败。直接用Python默认的UTF-8编码打开文件,遇到这类字节就会抛出UnicodeDecodeError。
解决方案
方案1:指定Latin-1编码打开文件
Latin-1编码能映射所有可能的字节值(0-255),不会出现解码错误,适合处理这类包含特殊字符的密码字典。修改文件打开代码:
with open("C:\\Users\\Shameer Kashif\\Desktop\\Python\\rockyouu.txt", "r", encoding="latin-1") as file:
方案2:二进制模式读取并手动处理
如果不确定编码,用二进制模式读取每行,再尝试解码(优先用UTF-8,失败则用Latin-1兜底):
with open("C:\\Users\\Shameer Kashif\\Desktop\\Python\\rockyouu.txt", "rb") as file: for line in file: attempt += 1 try: # 先尝试UTF-8解码 passwrd = line.strip().decode("utf-8") except UnicodeDecodeError: # 失败则用Latin-1兜底 passwrd = line.strip().decode("latin-1") hashed = hashlib.md5(passwrd.encode()).hexdigest() # 后续匹配逻辑保持不变
额外优化建议
- 统一哈希大小写:MD5哈希值通常是小写,将用户输入转小写可避免大小写不匹配:
user_hash = input("Enter your MD5 Hash to Decrypt: ").lower() - 细化异常捕获:把每行的解码逻辑放到内部try-except中,某一行出错不会中断整个破解流程,只会跳过该行继续处理:
with open("...", "r", encoding="latin-1") as file: for line in file: attempt += 1 try: passwrd = line.strip() hashed = hashlib.md5(passwrd.encode()).hexdigest() # 匹配逻辑... except UnicodeDecodeError: print(f"跳过第 {attempt} 行:无法解码的字符") continue - 线程安全改进:用
threading.Event()替代全局布尔变量loading_complete,多线程访问更安全:loading_complete = threading.Event() def show_loading_screen(): animation = "|/-\\|" idx = 0 while not loading_complete.is_set(): print("Cracking Possible Passwords... " + animation[idx % len(animation)], end="\r") idx += 1 time.sleep(0.1) # 标记加载完成时 loading_complete.set()
内容的提问来源于stack exchange,提问作者shameer kashif
相关产品推荐
相关产品推荐

