如何在Python运行时从文本文件创建自定义编解码器?
Great question! Let's figure out how to let users define their own custom encodings at runtime for your binary-to-text conversion script. Python's built-in codecs module is exactly what we need here—it lets us register custom encodings that work just like the standard ones you're used to (like cp1252). Here's a complete, practical approach:
Core Idea
Instead of hardcoding a modified cp1252 table, we'll let users provide their own 256-element decoding table (one entry per byte value 0-255). We'll then:
- Build a reverse mapping for encoding (characters back to bytes)
- Register this pair of encode/decode functions as a custom encoding
- Use the custom encoding just like any standard encoding in your script
Step-by-Step Implementation
1. Create a Function to Register Custom Encodings
This function will take a user-provided decoding table and register it as a named encoding with Python's codec system:
import codecs def register_custom_encoding(encoding_name, decoding_table): # Validate the decoding table is the correct length (256 entries for 0-255 bytes) if len(decoding_table) != 256: raise ValueError("Decoding table must contain exactly 256 characters (one for each byte value 0-255)") # Build the encoding table: map characters to their corresponding bytes # If multiple bytes map to the same character, we'll use the first byte encountered encoding_table = {} for byte_value, char in enumerate(decoding_table): if char not in encoding_table: encoding_table[char] = byte_value # Define the encoding function (str -> bytes) def encode(input_str, errors='strict'): output_bytes = [] for idx, char in enumerate(input_str): try: output_bytes.append(encoding_table[char]) except KeyError: if errors == 'strict': raise UnicodeEncodeError( encoding_name, char, idx, idx+1, f"Character '{char}' has no mapping in the custom encoding" ) elif errors == 'replace': output_bytes.append(0x3F) # Replace unknown chars with '?' elif errors == 'ignore': pass else: raise ValueError(f"Unsupported error handling mode: '{errors}'") return bytes(output_bytes), len(input_str) # Define the decoding function (bytes -> str) def decode(input_bytes, errors='strict'): output_str = [] for idx, byte in enumerate(input_bytes): try: output_str.append(decoding_table[byte]) except IndexError: # This shouldn't happen if we validated the table length, but just in case if errors == 'strict': raise UnicodeDecodeError( encoding_name, input_bytes, idx, idx+1, f"Byte 0x{byte:02X} is out of valid range (0-255)" ) elif errors == 'replace': output_str.append('\ufffd') # Standard replacement character elif errors == 'ignore': pass else: raise ValueError(f"Unsupported error handling mode: '{errors}'") return ''.join(output_str), len(input_bytes) # Register the codec with Python's codec system def codec_searcher(encoding): if encoding.lower() == encoding_name.lower(): return codecs.CodecInfo( encode=encode, decode=decode, name=encoding_name ) return None codecs.register(codec_searcher)
2. Use the Custom Encoding in Your Script
Once you've registered the encoding, you can use it just like any standard encoding for reading binary files and writing text:
# Example: User-provided custom decoding table (fill in all 256 entries as needed) user_decoding_table = ( '\x00', 'A', 'B', 'C', 'D', 'E', 'F', 'G', # Bytes 0-7 'H', 'I', 'J', 'K', 'L', 'M', 'N', 'O', # Bytes 8-15 # ... Add the remaining 240 characters here (one for each byte 16-255) ) # Register the encoding with a name of your choice (e.g., "user_custom_encoding") register_custom_encoding("user_custom_encoding", user_decoding_table) # Your existing workflow, now using the custom encoding def binary_to_text(input_bin_path, output_txt_path): # Read binary data with open(input_bin_path, 'rb') as bin_file: binary_data = bin_file.read() # Decode using the custom encoding text_data = binary_data.decode("user_custom_encoding", errors='replace') # Write to text file (use utf-8 to ensure all characters are saved correctly) with open(output_txt_path, 'w', encoding='utf-8') as text_file: text_file.write(text_data) # Run the conversion binary_to_text("input.bin", "output.txt") # Bonus: Convert text back to binary using the same custom encoding def text_to_binary(input_txt_path, output_bin_path): with open(input_txt_path, 'r', encoding='utf-8') as text_file: text_data = text_file.read() binary_data = text_data.encode("user_custom_encoding", errors='replace') with open(output_bin_path, 'wb') as bin_file: bin_file.write(binary_data) text_to_binary("input.txt", "output.bin")
Key Details to Note
- Error Handling: The implementation supports
strict,replace, andignoreerror modes—matching Python's standard encoding behavior. You can extend this if you need custom error handling. - Duplicate Characters: If multiple bytes map to the same character, the encoding function will use the first byte encountered in the decoding table. This ensures deterministic encoding.
- Validation: We check that the decoding table has exactly 256 entries to avoid index errors during decoding.
- Reusability: Once registered, the custom encoding can be used anywhere in your program that accepts an encoding name (like
open(),str.encode(), orbytes.decode()).
内容的提问来源于stack exchange,提问作者theflyingzamboni

