自定义Python SHA1实现与内置hashlib SHA1结果不一致问题排查
自定义SHA1实现与内置hashlib结果不符的修复方案
尝试用Python重实现SHA1算法,但自定义函数结果与hashlib.sha1不一致。已修正比特长度追加的问题,但结果仍有差异,代码如下:
from hashlib import sha1 as builtin_sha1 def rotl32(value: int, count: int) -> int: return ((value << count) | (value >> (32 - count))) & 0xffffffff def default_sha1(data: bytes) -> bytes: return builtin_sha1(data).digest() def sha1(data: bytes) -> bytes: # initialize variables h0 = 0x67452301 h1 = 0xefcdab89 h2 = 0x98badcfe h3 = 0x10325476 h4 = 0xc3d2e1f0 msg_len = len(data) # append 0x80 data += b"\x80" # append 0x00 until msg_len % 64 == 56 data += b"\x00" * ((56 - msg_len % 64) % 64) # append bit length as 64-bit big-endian integer data += (msg_len * 8).to_bytes(8, "big") # get the new length (now a multiple of 64) msg_len = len(data) for i in range(0, msg_len, 64): # for each chunk of 64 bytes # break the chunk into sixteen 32-bit big-endian words words = [int.from_bytes(data[i + j:i + j + 4], "big") for j in range(0, 64, 4)] # extend the sixteen 32-bit words into eighty 32-bit words for j in range(16, 80): words.append( rotl32((words[j - 3] ^ words[j - 8] ^ words[j - 14] ^ words[j - 16]), 1) ) # initialize hash value for this chunk a = h0 b = h1 c = h2 d = h3 e = h4 for j in range(80): if 0 <= j <= 19: f = (b & c) | ((~b) & d) k = 0x5a827999 elif 20 <= j <= 39: f = b ^ c ^ d k = 0x6ed9eba1 elif 40 <= j <= 59: f = (b & c) | (b & d) | (c & d) k = 0x8f1bbcdc else: # 60 <= j <= 79: f = b ^ c ^ d k = 0xca62c1d6 temp = (rotl32(a, 5) + f + e + k + words[j]) & 0xffffffff e = d d = c c = rotl32(b, 30) b = a a = temp # add this chunk's hash to result so far h0 = (h0 + a) & 0xffffffff h1 = (h1 + b) & 0xffffffff h2 = (h2 + c) & 0xffffffff h3 = (h3 + d) & 0xffffffff h4 = (h4 + e) & 0xffffffff # produce the final hash value return ((h0 << 128) | (h1 << 96) | (h2 << 64) | (h3 << 32) | h4).to_bytes(20, "big") if __name__ == "__main__": assert(sha1(b"hello") == default_sha1(b"hello")) # diff
问题定位
核心错误在消息填充阶段的0字节数量计算:
在追加0x80后,已经占用了1个字节,此时需要计算的是原始长度+1对64取模后的偏移,而非直接使用原始长度取模。原代码未考虑这1字节的占用,导致填充的0数量错误,破坏了消息块的对齐要求,最终导致哈希结果偏差。
修复后的代码
仅需修改填充0的数量计算逻辑:
from hashlib import sha1 as builtin_sha1 def rotl32(value: int, count: int) -> int: return ((value << count) | (value >> (32 - count))) & 0xffffffff def default_sha1(data: bytes) -> bytes: return builtin_sha1(data).digest() def sha1(data: bytes) -> bytes: # initialize variables h0 = 0x67452301 h1 = 0xefcdab89 h2 = 0x98badcfe h3 = 0x10325476 h4 = 0xc3d2e1f0 msg_len = len(data) # append 0x80 data += b"\x80" # 修复:计算填充0的数量时,考虑已追加的1字节 data += b"\x00" * ((56 - (msg_len + 1) % 64) % 64) # append bit length as 64-bit big-endian integer data += (msg_len * 8).to_bytes(8, "big") # get the new length (now a multiple of 64) msg_len = len(data) for i in range(0, msg_len, 64): # for each chunk of 64 bytes # break the chunk into sixteen 32-bit big-endian words words = [int.from_bytes(data[i + j:i + j + 4], "big") for j in range(0, 64, 4)] # extend the sixteen 32-bit words into eighty 32-bit words for j in range(16, 80): words.append( rotl32((words[j - 3] ^ words[j - 8] ^ words[j - 14] ^ words[j - 16]), 1) ) # initialize hash value for this chunk a = h0 b = h1 c = h2 d = h3 e = h4 for j in range(80): if 0 <= j <= 19: f = (b & c) | ((~b) & d) k = 0x5a827999 elif 20 <= j <= 39: f = b ^ c ^ d k = 0x6ed9eba1 elif 40 <= j <= 59: f = (b & c) | (b & d) | (c & d) k = 0x8f1bbcdc else: # 60 <= j <= 79: f = b ^ c ^ d k = 0xca62c1d6 temp = (rotl32(a, 5) + f + e + k + words[j]) & 0xffffffff e = d d = c c = rotl32(b, 30) b = a a = temp # add this chunk's hash to result so far h0 = (h0 + a) & 0xffffffff h1 = (h1 + b) & 0xffffffff h2 = (h2 + c) & 0xffffffff h3 = (h3 + d) & 0xffffffff h4 = (h4 + e) & 0xffffffff # produce the final hash value return ((h0 << 128) | (h1 << 96) | (h2 << 64) | (h3 << 32) | h4).to_bytes(20, "big") if __name__ == "__main__": assert(sha1(b"hello") == default_sha1(b"hello")) # 现在断言通过
验证
修复后,运行代码会发现断言通过,自定义SHA1实现的结果与hashlib.sha1完全一致。
内容的提问来源于stack exchange,提问作者Fayeure
相关产品推荐
相关产品推荐

