霍夫曼算法:0/1字符串转二进制后如何转回原字符串?
Got it, let's work through reversing your toBytes function properly. The key problem here is that when your original 0/1 string isn't a multiple of 8 bits long, converting the bytes back directly will add extra leading zeros to the last chunk—because each byte always translates to 8 bits, even if the original chunk was shorter. To fix this, we need to track the original length of the bit string, either separately or by embedding it in the byte data.
Why Your Current Approach Fails
Let's use your example:
- Input bit string:
"01111101011000"(14 bits total) - Your
toBytessplits it into two chunks: first 8 bits (01111101→125→}) and last 6 bits (011000→24→\x18), resulting inb'}\x18'. - If you just convert each byte back to binary, you get
01111101+00011000="0111110100011000"—which has 2 extra leading zeros in the second chunk, not matching the original string.
Solution 1: Track Original Bit Length Separately
If you can pass the original length of the bit string along with the bytes, you can write a decoder that truncates the full 8-bit chunks to the correct length:
def fromBytes(b, original_length): bits = [] for byte in b: # Convert each byte to an 8-bit string, padding with leading zeros if needed bits.append(f"{byte:08b}") # Join all 8-bit chunks into one long string full_bit_string = ''.join(bits) # Truncate to the original length to remove extra leading zeros from the last chunk return full_bit_string[:original_length]
Test This With Your Example
encoded_bytes = b'}\x18' original_bit_string = "01111101011000" decoded = fromBytes(encoded_bytes, len(original_bit_string)) print(decoded) # Output: "01111101011000" (matches original)
Solution 2: Embed the Length in the Byte Stream (More Practical)
For real-world use (like saving to a file or sending over a network), you'll want to embed the original bit length directly into the byte data. We can use Python's struct module to pack the length as a 4-byte integer at the start of the byte stream:
Updated Encoder with Length
import struct def toBytesWithLength(data): # Pack the original bit length as a 4-byte big-endian integer length_bytes = struct.pack('>I', len(data)) byte_arr = bytearray() for i in range(0, len(data), 8): chunk = data[i:i+8] byte_arr.append(int(chunk, 2)) # Return length bytes + encoded data bytes return length_bytes + bytes(byte_arr)
Corresponding Decoder
def fromBytesWithLength(b): # Unpack the first 4 bytes to get the original bit length original_length = struct.unpack('>I', b[:4])[0] # Process the remaining bytes data_bytes = b[4:] bits = [] for byte in data_bytes: bits.append(f"{byte:08b}") full_bit_string = ''.join(bits) return full_bit_string[:original_length]
Test This
original = "01111101011000" encoded = toBytesWithLength(original) decoded = fromBytesWithLength(encoded) print(decoded == original) # Output: True
Key Notes
- Always use
f"{byte:08b}"to convert bytes to 8-bit strings—avoidbin(byte)[2:], since it doesn't pad leading zeros (e.g.,bin(24)gives'11000'instead of'00011000'). - Embedding the length is better for persistence/transmission because you don't have to track the length separately.
内容的提问来源于stack exchange,提问作者OldLazarus

