Python中如何将任意长度比特串写入二进制文件并正确读取
Hey there! Let's sort out this bit storage problem for you. The issue with your current approach is that you're treating each individual '0' or '1' as a full byte, which is why your 194-bit string ends up taking 194 bytes of space—total overkill! We need to pack those bits into actual bytes (8 bits per byte) to save space, plus handle the leftover bits at the end.
Why Your Current Code Isn't Working
When you do bytes(list(map(int, bitstring))), each '0' becomes the byte 0x00 (binary 00000000) and each '1' becomes 0x01 (binary 00000001). That means every single bit in your original string is being stored as an entire 8-bit byte—hence the 194-byte file size.
The Fix: Pack Bits Into Bytes
To get the minimal file size, we need to:
- Pad the bitstring to make its length a multiple of 8 (since files are stored in whole bytes).
- Convert groups of 8 bits into single bytes.
- Store the original length of the bitstring, so we can trim off the padding when reading back.
Here's a complete implementation:
Writing the Bitstring
import struct bitstring = '10110101111111001101101010011011111010011010110001010101011100010110100010001001110001110100011111010001100011011110010100110000010111101011001011010111111100000110000000001001101000010110000111' filename = 'bits.bin' with open(filename, 'wb') as f: # Store the original length (use 'B' for 1-byte storage if length < 256) original_length = len(bitstring) f.write(struct.pack('B', original_length)) # 'B' = unsigned char (1 byte) # Calculate padding needed to make length a multiple of 8 padding = (8 - original_length % 8) % 8 padded_bitstring = bitstring + '0' * padding # Convert padded string to bytes byte_array = bytearray() for i in range(0, len(padded_bitstring), 8): # Convert 8-bit chunk to integer, then to byte byte = int(padded_bitstring[i:i+8], 2) byte_array.append(byte) f.write(byte_array)
Reading the Bitstring Back
import struct filename = 'bits.bin' with open(filename, 'rb') as f: # Read the original length first original_length = struct.unpack('B', f.read(1))[0] # Read the rest of the bytes byte_data = f.read() # Convert bytes back to a full bitstring recovered_bitstring = '' for byte in byte_data: # Convert each byte to an 8-bit string (preserve leading zeros!) recovered_bitstring += format(byte, '08b') # Trim off the padding to get the original string recovered_bitstring = recovered_bitstring[:original_length] # Verify it's correct print(f"Original length: {len(bitstring)} | Recovered length: {len(recovered_bitstring)}") print(f"First 30 bits match: {bitstring[:30] == recovered_bitstring[:30]}")
What This Does
- For your 194-bit string, we add 2 padding zeros to make it 196 bits (24 full bytes + 1 partial byte = 25 bytes total for the bit data).
- We store the original length in 1 byte (since 194 < 256), so the total file size is 1 + 25 = 26 bytes—almost exactly the 25 you wanted (we can't avoid storing the original length, otherwise we wouldn't know how many bits to keep from the last byte).
- When reading, we convert each byte back to an 8-bit string, then trim to the original length to discard the padding.
Key Notes
- Using
struct.pack('B')works for bitstrings up to 255 bits. If you need longer strings, switch to'H'(2 bytes) or'I'(4 bytes) for the length storage. - The
format(byte, '08b')is crucial—it ensures each byte is converted to an 8-bit string with leading zeros, so we don't lose bits from smaller numbers (e.g., the byte5becomes00000101, not just101).
内容的提问来源于stack exchange,提问作者Nip

