Python生成BIP-39助记词验证失败问题求助
BIP-39助记词生成无效的问题排查与修复
问题描述
我尝试用Python生成符合BIP-39标准的比特币钱包助记词,但生成的助记词无法通过校验工具验证。严格遵循BIP-39规范实现,但第24个校验位对应的助记词始终导致整个助记词无效,参考过其他代码仍未解决,求错误分析和修复方案。
原代码
from hashlib import sha256 import secrets #following the instructions here: https://github.com/bitcoin/bips/blob/master/bip-0039.mediawiki #depending on the number of words, we take the value for ENT, and CS.a word_number=24 size_ENT=256 size_CS=int(size_ENT/32) with open("Bip39-wordlist.txt", "r") as wordlist_file: words = [word.strip() for word in wordlist_file.readlines()] #First, an initial entropy of ENT bits is generated. n_bytes=int(size_ENT/8) random_bytes = secrets.token_bytes(n_bytes) random_bits = ''.join(['{:08b}'.format(b) for b in random_bytes]) INITIAL_ENTROPY = random_bits[:size_ENT] assert(len(INITIAL_ENTROPY)==size_ENT) encoded=INITIAL_ENTROPY.encode('utf-8') hash=sha256(encoded).digest() bhash=''.join(format(byte, '08b') for byte in hash) assert(len(bhash)==256) #the first ENT / 32 bits of its SHA256 hash CS=bhash[:size_CS] #This checksum is appended to the end of the initial entropy. FINAL_ENTROPY=INITIAL_ENTROPY+CS assert(len(FINAL_ENTROPY)==size_ENT+size_CS) #Next, these concatenated bits are split into groups of 11 bits, # each encoding a number from 0-2047, serving as an index into a wordlist. for t in range(word_number): #split into groups of 11 bits, extracted_bits=FINAL_ENTROPY[11*t:11*(t+1)] # each encoding a number from 0-2047, word_index=int(extracted_bits,2) #serving as an index into a wordlist. if t==0: words_extracted= words[word_index] else: words_extracted+=' '+words[word_index] print (words_extracted)
无效助记词示例
- kitten oak breeze dismiss breeze reduce stem symbol trend input thunder old burden brisk level hard luggage alarm upper creek deputy desert diesel primary
- wave flee narrow notable budget hamster layer potato menu security wall shove save mobile badge nephew blouse major cute park margin entry drink mask
错误原因分析
核心问题出在校验位计算的哈希源错误:
你将初始熵的二进制字符串(如'101010...')编码为UTF-8后计算SHA256哈希,但BIP-39规范要求直接对初始熵的原始字节数据计算哈希,而非二进制字符串的UTF-8编码。
举个简单例子:初始熵中的一个字节是0xAB,对应的二进制字符串是'10101011'。你把这个字符串转成UTF-8字节(每个字符对应一个字节,比如'1'是0x31,'0'是0x30)再哈希,这和直接对0xAB这个原始字节哈希的结果完全不同,导致校验位计算错误,最终助记词无法通过校验。
修复方案
直接使用生成的random_bytes(原始熵字节)计算SHA256哈希,无需转成二进制字符串再编码。修复后的完整代码如下:
from hashlib import sha256 import secrets # 遵循BIP-39规范 word_number = 24 size_ENT = 256 size_CS = int(size_ENT / 32) # 加载BIP-39词表 with open("Bip39-wordlist.txt", "r") as wordlist_file: words = [word.strip() for word in wordlist_file.readlines()] # 生成初始熵(ENT位) n_bytes = int(size_ENT / 8) random_bytes = secrets.token_bytes(n_bytes) # 将原始字节转成二进制字符串,用于后续拼接校验位 random_bits = ''.join(['{:08b}'.format(b) for b in random_bytes]) INITIAL_ENTROPY = random_bits[:size_ENT] assert(len(INITIAL_ENTROPY) == size_ENT) # 关键修复:直接对原始熵字节计算SHA256哈希 hash = sha256(random_bytes).digest() bhash = ''.join(format(byte, '08b') for byte in hash) assert(len(bhash) == 256) # 取哈希的前CS位作为校验位 CS = bhash[:size_CS] # 拼接初始熵和校验位 FINAL_ENTROPY = INITIAL_ENTROPY + CS assert(len(FINAL_ENTROPY) == size_ENT + size_CS) # 分割成11位一组,映射到词表 words_extracted = [] for t in range(word_number): extracted_bits = FINAL_ENTROPY[11*t:11*(t+1)] word_index = int(extracted_bits, 2) words_extracted.append(words[word_index]) print(' '.join(words_extracted))
额外优化说明
- 改用列表
append后用' '.join()拼接助记词,比字符串累加更高效易读。 - 保留所有断言,确保每一步的长度符合BIP-39规范,便于调试。
内容的提问来源于stack exchange,提问作者Pietro Speroni
相关产品推荐
相关产品推荐

