含Unicode字符的消息哈希:NodeJS与Python结果不一致的解决方法
问题描述
我需要对包含特殊字节(如\x1a、\xab)的消息进行SHA256哈希运算,但Node.js与Python的计算结果不一致。期望Node.js能输出和Python相同的哈希值:947270b0b8041a92ba82ef37661b692a4a150532b88de59bf95e965ceb5c07f8。
现有代码
Node.js(使用CryptoJS)
const CryptoJS = require('crypto-js'); let message = '\x1aSmartCash Signed Message:\n\xabCTxIn(COutPoint(7bb8ad134928a003752beb098471af5a66fc5475ff96b5ba4c2e1c4cbac3aa13, 0), scriptSig=)000000000002c5c2ef4afc588492773e6bbb18e4f2374b6dc159ef257667bd881667906410'; let hash = CryptoJS.SHA256(message).toString(); console.log('hash', hash); // 989c004534c6962293c95a9438bdb926c92c1d8b4dec0f4f1e535defa171e5fe
Python
import hashlib def to_bytes(something, encoding='utf8'): """ cast string to bytes() like object, but for python2 support it's bytearray copy """ if isinstance(something, bytes): return something if isinstance(something, str): return something.encode(encoding) elif isinstance(something, bytearray): return bytes(something) else: raise TypeError("Not a string or bytes like object") def sha256(x): x = to_bytes(x, 'utf8') return bytes(hashlib.sha256(x).digest()) def Hash_Sha256(x): x = to_bytes(x, 'utf8') out = bytes(sha256(x)) return out message = b'\x1aSmartCash Signed Message:\n\xabCTxIn(COutPoint(7bb8ad134928a003752beb098471af5a66fc5475ff96b5ba4c2e1c4cbac3aa13, 0), scriptSig=)000000000002c5c2ef4afc588492773e6bbb18e4f2374b6dc159ef257667bd881667906410' print('hash', Hash_Sha256(message).hex()) # 947270b0b8041a92ba82ef37661b692a4a150532b88de59bf95e965ceb5c07f8
问题根源
- Python里的
b'\x1a'、b'\xab'是单字节原始数据,直接对应0x1A、0xAB的字节值,没有编码转换。 - Node.js中,字符串
'\x1a'、'\xab'被CryptoJS默认按UTF-8编码处理:其中\xab不属于合法UTF-8字符,会被转译为UTF-8的替代字符序列(0xEF 0xBF 0xBD),导致最终参与哈希计算的字节序列和Python完全不同。
解决方案
要让两边使用完全一致的字节序列计算哈希,需要让CryptoJS按单字节编码规则处理字符串,或者直接传入原始字节数组:
方案1:用CryptoJS的enc.Latin1编码字符串
Latin1(ISO-8859-1)编码会将每个字符映射为单字节,和Python的bytes行为完全匹配:
const CryptoJS = require('crypto-js'); let message = '\x1aSmartCash Signed Message:\n\xabCTxIn(COutPoint(7bb8ad134928a003752beb098471af5a66fc5475ff96b5ba4c2e1c4cbac3aa13, 0), scriptSig=)000000000002c5c2ef4afc588492773e6bbb18e4f2374b6dc159ef257667bd881667906410'; // 将字符串按Latin1解析为字节数组后计算哈希 let hash = CryptoJS.SHA256(CryptoJS.enc.Latin1.parse(message)).toString(); console.log('hash', hash); // 947270b0b8041a92ba82ef37661b692a4a150532b88de59bf95e965ceb5c07f8
方案2:直接用Node.js Buffer传递原始字节
如果消息本身是原始字节数据,用Buffer定义更直观,和Python的b''语法逻辑一致:
const CryptoJS = require('crypto-js'); const { Buffer } = require('buffer'); // 用Buffer拼接原始字节,和Python的字节序列完全对齐 const messageBuffer = Buffer.concat([ Buffer.from([0x1a]), Buffer.from('SmartCash Signed Message:\n'), Buffer.from([0xab]), Buffer.from('CTxIn(COutPoint(7bb8ad134928a003752beb098471af5a66fc5475ff96b5ba4c2e1c4cbac3aa13, 0), scriptSig=)000000000002c5c2ef4afc588492773e6bbb18e4f2374b6dc159ef257667bd881667906410') ]); // 将Buffer转换为CryptoJS可处理的WordArray let hash = CryptoJS.SHA256(CryptoJS.lib.WordArray.create(messageBuffer)).toString(); console.log('hash', hash); // 947270b0b8041a92ba82ef37661b692a4a150532b88de59bf95e965ceb5c07f8
验证
两种方案都能输出和Python完全一致的哈希值,解决了特殊字节编码不一致的问题。
内容的提问来源于stack exchange,提问作者CODEHUB
相关产品推荐
相关产品推荐

