You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

霍夫曼算法:0/1字符串转二进制后如何转回原字符串?

Reverse Huffman Bit String from Bytes

Got it, let's work through reversing your toBytes function properly. The key problem here is that when your original 0/1 string isn't a multiple of 8 bits long, converting the bytes back directly will add extra leading zeros to the last chunk—because each byte always translates to 8 bits, even if the original chunk was shorter. To fix this, we need to track the original length of the bit string, either separately or by embedding it in the byte data.

Why Your Current Approach Fails

Let's use your example:

  • Input bit string: "01111101011000" (14 bits total)
  • Your toBytes splits it into two chunks: first 8 bits (01111101 → 125 → }) and last 6 bits (011000 → 24 → \x18), resulting in b'}\x18'.
  • If you just convert each byte back to binary, you get 01111101 + 00011000 = "0111110100011000"—which has 2 extra leading zeros in the second chunk, not matching the original string.

Solution 1: Track Original Bit Length Separately

If you can pass the original length of the bit string along with the bytes, you can write a decoder that truncates the full 8-bit chunks to the correct length:

def fromBytes(b, original_length):
    bits = []
    for byte in b:
        # Convert each byte to an 8-bit string, padding with leading zeros if needed
        bits.append(f"{byte:08b}")
    # Join all 8-bit chunks into one long string
    full_bit_string = ''.join(bits)
    # Truncate to the original length to remove extra leading zeros from the last chunk
    return full_bit_string[:original_length]

Test This With Your Example

encoded_bytes = b'}\x18'
original_bit_string = "01111101011000"
decoded = fromBytes(encoded_bytes, len(original_bit_string))
print(decoded)  # Output: "01111101011000" (matches original)

Solution 2: Embed the Length in the Byte Stream (More Practical)

For real-world use (like saving to a file or sending over a network), you'll want to embed the original bit length directly into the byte data. We can use Python's struct module to pack the length as a 4-byte integer at the start of the byte stream:

Updated Encoder with Length

import struct

def toBytesWithLength(data):
    # Pack the original bit length as a 4-byte big-endian integer
    length_bytes = struct.pack('>I', len(data))
    byte_arr = bytearray()
    for i in range(0, len(data), 8):
        chunk = data[i:i+8]
        byte_arr.append(int(chunk, 2))
    # Return length bytes + encoded data bytes
    return length_bytes + bytes(byte_arr)

Corresponding Decoder

def fromBytesWithLength(b):
    # Unpack the first 4 bytes to get the original bit length
    original_length = struct.unpack('>I', b[:4])[0]
    # Process the remaining bytes
    data_bytes = b[4:]
    bits = []
    for byte in data_bytes:
        bits.append(f"{byte:08b}")
    full_bit_string = ''.join(bits)
    return full_bit_string[:original_length]

Test This

original = "01111101011000"
encoded = toBytesWithLength(original)
decoded = fromBytesWithLength(encoded)
print(decoded == original)  # Output: True

Key Notes

  • Always use f"{byte:08b}" to convert bytes to 8-bit strings—avoid bin(byte)[2:], since it doesn't pad leading zeros (e.g., bin(24) gives '11000' instead of '00011000').
  • Embedding the length is better for persistence/transmission because you don't have to track the length separately.

内容的提问来源于stack exchange,提问作者OldLazarus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:11:34