You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python运行时从文本文件创建自定义编解码器?

Great question! Let's figure out how to let users define their own custom encodings at runtime for your binary-to-text conversion script. Python's built-in codecs module is exactly what we need here—it lets us register custom encodings that work just like the standard ones you're used to (like cp1252). Here's a complete, practical approach:

Core Idea

Instead of hardcoding a modified cp1252 table, we'll let users provide their own 256-element decoding table (one entry per byte value 0-255). We'll then:

  1. Build a reverse mapping for encoding (characters back to bytes)
  2. Register this pair of encode/decode functions as a custom encoding
  3. Use the custom encoding just like any standard encoding in your script

Step-by-Step Implementation

1. Create a Function to Register Custom Encodings

This function will take a user-provided decoding table and register it as a named encoding with Python's codec system:

import codecs

def register_custom_encoding(encoding_name, decoding_table):
    # Validate the decoding table is the correct length (256 entries for 0-255 bytes)
    if len(decoding_table) != 256:
        raise ValueError("Decoding table must contain exactly 256 characters (one for each byte value 0-255)")
    
    # Build the encoding table: map characters to their corresponding bytes
    # If multiple bytes map to the same character, we'll use the first byte encountered
    encoding_table = {}
    for byte_value, char in enumerate(decoding_table):
        if char not in encoding_table:
            encoding_table[char] = byte_value
    
    # Define the encoding function (str -> bytes)
    def encode(input_str, errors='strict'):
        output_bytes = []
        for idx, char in enumerate(input_str):
            try:
                output_bytes.append(encoding_table[char])
            except KeyError:
                if errors == 'strict':
                    raise UnicodeEncodeError(
                        encoding_name, char, idx, idx+1,
                        f"Character '{char}' has no mapping in the custom encoding"
                    )
                elif errors == 'replace':
                    output_bytes.append(0x3F)  # Replace unknown chars with '?'
                elif errors == 'ignore':
                    pass
                else:
                    raise ValueError(f"Unsupported error handling mode: '{errors}'")
        return bytes(output_bytes), len(input_str)
    
    # Define the decoding function (bytes -> str)
    def decode(input_bytes, errors='strict'):
        output_str = []
        for idx, byte in enumerate(input_bytes):
            try:
                output_str.append(decoding_table[byte])
            except IndexError:
                # This shouldn't happen if we validated the table length, but just in case
                if errors == 'strict':
                    raise UnicodeDecodeError(
                        encoding_name, input_bytes, idx, idx+1,
                        f"Byte 0x{byte:02X} is out of valid range (0-255)"
                    )
                elif errors == 'replace':
                    output_str.append('\ufffd')  # Standard replacement character
                elif errors == 'ignore':
                    pass
                else:
                    raise ValueError(f"Unsupported error handling mode: '{errors}'")
        return ''.join(output_str), len(input_bytes)
    
    # Register the codec with Python's codec system
    def codec_searcher(encoding):
        if encoding.lower() == encoding_name.lower():
            return codecs.CodecInfo(
                encode=encode,
                decode=decode,
                name=encoding_name
            )
        return None
    
    codecs.register(codec_searcher)

2. Use the Custom Encoding in Your Script

Once you've registered the encoding, you can use it just like any standard encoding for reading binary files and writing text:

# Example: User-provided custom decoding table (fill in all 256 entries as needed)
user_decoding_table = (
    '\x00', 'A', 'B', 'C', 'D', 'E', 'F', 'G',  # Bytes 0-7
    'H', 'I', 'J', 'K', 'L', 'M', 'N', 'O',     # Bytes 8-15
    # ... Add the remaining 240 characters here (one for each byte 16-255)
)

# Register the encoding with a name of your choice (e.g., "user_custom_encoding")
register_custom_encoding("user_custom_encoding", user_decoding_table)

# Your existing workflow, now using the custom encoding
def binary_to_text(input_bin_path, output_txt_path):
    # Read binary data
    with open(input_bin_path, 'rb') as bin_file:
        binary_data = bin_file.read()
    
    # Decode using the custom encoding
    text_data = binary_data.decode("user_custom_encoding", errors='replace')
    
    # Write to text file (use utf-8 to ensure all characters are saved correctly)
    with open(output_txt_path, 'w', encoding='utf-8') as text_file:
        text_file.write(text_data)

# Run the conversion
binary_to_text("input.bin", "output.txt")

# Bonus: Convert text back to binary using the same custom encoding
def text_to_binary(input_txt_path, output_bin_path):
    with open(input_txt_path, 'r', encoding='utf-8') as text_file:
        text_data = text_file.read()
    
    binary_data = text_data.encode("user_custom_encoding", errors='replace')
    
    with open(output_bin_path, 'wb') as bin_file:
        bin_file.write(binary_data)

text_to_binary("input.txt", "output.bin")

Key Details to Note

  • Error Handling: The implementation supports strict, replace, and ignore error modes—matching Python's standard encoding behavior. You can extend this if you need custom error handling.
  • Duplicate Characters: If multiple bytes map to the same character, the encoding function will use the first byte encountered in the decoding table. This ensures deterministic encoding.
  • Validation: We check that the decoding table has exactly 256 entries to avoid index errors during decoding.
  • Reusability: Once registered, the custom encoding can be used anywhere in your program that accepts an encoding name (like open(), str.encode(), or bytes.decode()).

内容的提问来源于stack exchange,提问作者theflyingzamboni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:22:33