You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算结构体的keccak256/sha256哈希?为何sha256(a)≠sha256(b)及实现原理?

Hey there! Let's tackle your two questions about cryptographic hashing—first calculating hashes for structs, then diving into SHA256's implementation and why your hashes don't match.

Calculating Keccak256/SHA256 Hashes for Structs

The core idea here is that hash functions like Keccak256 and SHA256 operate on raw byte streams. To hash a struct, you first need to serialize all its fields into a consistent, fixed-format byte sequence—the exact method depends on the programming language you're using.

Example 1: Solidity (Common for Blockchain Use Cases)

Solidity has built-in utilities to handle struct serialization. You'll use abi.encode or abi.encodePacked to convert struct fields into bytes, then pass that to the hash function:

// Define your struct
struct User {
    uint256 userId;
    string username;
    address wallet;
}

function hashUser(User memory user) public pure returns (bytes32) {
    // Option 1: abi.encode (safer, handles type alignment to avoid collisions)
    return keccak256(abi.encode(user.userId, user.username, user.wallet));
    
    // Option 2: abi.encodePacked (more compact, but risk of hash collisions if fields overlap in byte format)
    // return keccak256(abi.encodePacked(user.userId, user.username, user.wallet));
}

Note: Always prefer abi.encode unless you specifically need the compactness of encodePacked—the latter can lead to collisions when combining fields of variable length (e.g., two different string pairs that pack to the same bytes).

Example 2: Python (General-Purpose)

For Python, you'll need to manually serialize each struct field to bytes (ensuring consistent byte order and length), then hash the combined byte stream. We'll use the pycryptodome library for hashing:

from Crypto.Hash import SHA256, keccak
from struct import pack

def hash_user_struct(user_id: int, username: str, wallet: str) -> tuple[str, str]:
    # Serialize each field to fixed-format bytes
    user_id_bytes = pack('>I', user_id)  # Big-endian 4-byte unsigned int
    username_bytes = username.encode('utf-8').ljust(32, b'\x00')  # Pad to 32 bytes
    wallet_bytes = bytes.fromhex(wallet.strip('0x'))  # Convert Ethereum address to 20-byte raw bytes
    
    # Combine all bytes into a single stream
    raw_data = user_id_bytes + username_bytes + wallet_bytes
    
    # Calculate hashes
    sha256_result = SHA256.new(raw_data).hexdigest()
    keccak256_result = keccak.new(digest_bits=256, data=raw_data).hexdigest()
    
    return sha256_result, keccak256_result

Critical: If you're hashing structs across different languages, make sure all serialization rules (byte order, field length, padding) are identical—even a tiny difference (e.g., little-endian vs big-endian) will produce a completely different hash.

SHA256 Implementation & Why sha256(a) != sha256(b)

First, let's break down how SHA256 works at a high level:

Core SHA256 Implementation Steps

SHA256 is a deterministic cryptographic hash function that takes arbitrary input and produces a fixed 256-bit output. It follows these steps:

  1. Preprocessing: Pad the input so its total length is a multiple of 512 bits. The last 64 bits store the original input length in binary.
  2. Initialize Hash Constants: Start with 8 32-bit constants derived from the square roots of the first 8 prime numbers.
  3. Process 512-Bit Blocks:
    • Split each 512-bit block into 16 32-bit words, then expand them into 64 32-bit words using bitwise operations.
    • Run a 64-round compression loop using 64 additional constants (from cube roots of the first 64 primes) and logical functions (Ch, Maj, Σ0, Σ1, σ0, σ1) to update the hash values.
  4. Finalize: Combine the 8 updated hash values into a single 256-bit string.

Why Your Hashes Don't Match

SHA256 is completely deterministic—if sha256(a) != sha256(b), the input byte streams for a and b are definitely different. Common culprits include:

  • Hidden Input Differences: Invisible characters (newlines, spaces, tabs), different string encodings (UTF-8 vs GBK), or trailing null bytes.
  • Inconsistent Serialization: For structs, mismatched byte order, field padding, or field order (e.g., hashing (id, name) vs (name, id) will produce different results).
  • Type Mismatches: Treating a number as an integer vs a string (e.g., the integer 10 is 0x0A in bytes, while the string "10" is 0x3130).
  • Encoding Errors: For text inputs, forgetting to encode strings to bytes (e.g., passing a Python string directly to a hash function instead of using .encode('utf-8')).

Quick example of how tiny differences change the hash:

from Crypto.Hash import SHA256

a = "hello"
b = "hello "  # Notice the trailing space
print(SHA256.new(a.encode()).hexdigest())  # Output: 2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824
print(SHA256.new(b.encode()).hexdigest())  # Output: 926e2161640be84e9046f980545d40f17f460f7149ac67689780f4fdb703044c

内容的提问来源于stack exchange,提问作者user2060883

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:05:34