如何让hashlib代码同时兼容Python2与Python3
问题描述
这段代码在Python2环境运行正常,但切换到Python3环境时会抛出错误:
Checksum error: bob.tgz, Unicode-objects must be encoded before Hashing
尝试调整sha256对象使用字节对象后,又出现新的解码错误:
Checksum error: bob.tgz, 'utf-8' codec can't decode byte 0x8b in Position 1: invalid start byte
原代码如下:
import hashlib filepath = "bob.tgz" filepath_hashed = "bob.tgz.hashed" thekey = "asdf1234" try: m = hashlib.sha256() file_contents = None with open(filepath, "r") as f_in: file_contents = f_in.read() m.update(file_contents) m.update(thekey) with open(filepath_hashed, "w") as f_out: f_out.write(file_contents) f_out.write("~~CHECKSUM~~%s" % m.hexdigest()) except Exception as e: print("Checksum error: %s, %s" % (filepath, e))
解决方案
问题本质是Python2和3对文本、二进制数据的处理逻辑不一致:
- Python2中
open()默认文本模式,但读取二进制文件时直接返回字节串;Python3中文本模式会自动按UTF-8解码字节,二进制文件的非UTF-8字节会触发解码错误。 - Python3的
hashlib.sha256.update()只接受字节对象,不支持字符串。
按以下三点修改即可兼容两个版本:
- 二进制模式读写文件:用
rb(读二进制)和wb(写二进制)打开文件,绕开编码/解码步骤,确保读取的内容是字节串。 - 密钥转字节串:把字符串类型的
thekey编码为字节后再传入update()。 - 校验和转字节写入:写入校验和字符串时,要转为字节再写入二进制文件。
修改后的兼容代码:
import hashlib filepath = "bob.tgz" filepath_hashed = "bob.tgz.hashed" thekey = "asdf1234" try: m = hashlib.sha256() file_contents = None # 二进制模式读取文件,避免编码错误 with open(filepath, "rb") as f_in: file_contents = f_in.read() m.update(file_contents) # 将字符串密钥编码为字节串 m.update(thekey.encode('utf-8')) # 二进制模式写入文件 with open(filepath_hashed, "wb") as f_out: f_out.write(file_contents) # 将校验和字符串转为字节后写入 checksum_str = "~~CHECKSUM~~%s" % m.hexdigest() f_out.write(checksum_str.encode('utf-8')) except Exception as e: print("Checksum error: %s, %s" % (filepath, str(e)))
内容的提问来源于stack exchange,提问作者user740521
相关产品推荐
相关产品推荐

