如何将由0和1组成的二进制文本文件还原为原始文件?
Hey there, let's break down how to solve this problem. You've got a text file filled with nothing but 0s and 1s—this is just a human-readable version of raw binary data. To get your original file back, we just need to reverse that conversion process. Here's how to do it with practical code examples:
Core Concept
Your text file takes every byte from the original file, converts each bit to a character ('0' or '1'), and writes those characters out. The reverse process works like this:
- Read all valid 0/1 characters (ignoring any extra newlines or whitespace that might be present)
- Group the characters into chunks of 8 (since each byte is 8 bits)
- Convert each 8-character binary string into a single byte
- Write all those bytes to a new file—that's your original file!
Python Implementation (Most Straightforward)
Python is great for this because it handles string manipulation and file I/O smoothly. Here's a ready-to-use function:
def binary_text_to_original(input_txt_path, output_bin_path): # Read input and filter out non-0/1 characters with open(input_txt_path, 'r') as input_file: binary_chars = ''.join(char for char in input_file.read() if char in ('0', '1')) # Handle cases where total bits aren't a multiple of 8 if len(binary_chars) % 8 != 0: pad_length = 8 - (len(binary_chars) % 8) binary_chars += '0' * pad_length print(f"Note: Input length wasn't a multiple of 8. Added {pad_length} trailing 0s.") # Convert 8-bit chunks to bytes byte_data = bytearray() for i in range(0, len(binary_chars), 8): byte = int(binary_chars[i:i+8], 2) byte_data.append(byte) # Write the final binary file with open(output_bin_path, 'wb') as output_file: output_file.write(byte_data) # Example usage: replace with your actual file paths binary_text_to_original('binary_input.txt', 'restored_file.bin')
For Large Files (Memory-Friendly Version)
If your text file is too big to load all at once, use this streaming approach to avoid memory issues:
def large_binary_text_to_original(input_txt_path, output_bin_path): with open(input_txt_path, 'r') as input_file, open(output_bin_path, 'wb') as output_file: char_buffer = '' for line in input_file: # Add valid 0/1 characters to the buffer char_buffer += ''.join(char for char in line if char in ('0', '1')) # Process full 8-bit chunks as we go while len(char_buffer) >= 8: byte = int(char_buffer[:8], 2) output_file.write(bytes([byte])) char_buffer = char_buffer[8:] # Handle leftover bits at the end if char_buffer: pad_length = 8 - len(char_buffer) char_buffer += '0' * pad_length byte = int(char_buffer, 2) output_file.write(bytes([byte])) print(f"Note: {len(char_buffer)-pad_length} leftover bits padded with {pad_length} 0s.")
Shell Script (Linux/macOS)
If you prefer command-line tools over writing a script, chain these commands together:
# Filter non-0/1 chars → split into 8-character lines → convert to bytes → write to file tr -cd '01' < binary_input.txt | fold -w8 | awk '{printf "%c", strtonum("0b"$0)}' > restored_file.bin
Important Notes
- Bit Order Check: The code above assumes the most significant bit (MSB) comes first (standard for most binary representations). If your file uses least significant bit (LSB) first, reverse each 8-character chunk before converting (e.g.,
binary_chars[i:i+8][::-1]in Python). - Padding: The code adds trailing 0s if the total bits aren't divisible by 8. If you know the original file shouldn't have padding, modify the code to throw an error instead.
内容的提问来源于stack exchange,提问作者Ggdev

