如何解码三进制及其他非常规进制字符串?附Python脚本求助
Hey there! Let's break down what's going wrong with your current approach and walk through reliable solutions for converting non-standard bases (like base-3, base-5, etc.) to plaintext.
What's Wrong With Your Current Code?
Your core issue is how you're mapping base-3 digits to binary: manually converting 0→00, 1→01, 2→10 creates inconsistent value translations. For example:
- Base-3
12should equal1*3 + 2 = 5(binary101), but your method turns it into01 + 10 = 0110(decimal6). This off-by-one error cascades through the entire string, leading to garbage output.
Solution 1: Convert the Entire Base String to a Decimal Integer First
This is the most universal approach for cases where the base-encoded data represents a single large integer (similar to how base64 works under the hood). Here's the fixed code:
base = 3 base_data = "" # Read the base-3 file content with open(f"./base{base}", "r") as b3: for line in b3: base_data += line.strip("\r\n") # Step 1: Convert the full base-3 string to a decimal integer decimal_num = int(base_data, base) # Step 2: Convert the decimal integer to a byte string # Calculate the number of bytes needed to store the integer byte_length = (decimal_num.bit_length() + 7) // 8 # Try 'big' first; if that fails, switch to 'little' for byte order byte_data = decimal_num.to_bytes(byte_length, byteorder="big") # Step 3: Decode bytes to plaintext (try different encodings if utf-8 fails) try: plaintext = byte_data.decode("utf-8") except UnicodeDecodeError: # Fallback to replace invalid characters, or try 'ascii'/'gbk' instead plaintext = byte_data.decode("utf-8", errors="replace") print(plaintext)
Solution 2: Fixed-Length Group Mapping (If Data Is Chunked)
If the base-3 data was encoded by grouping digits to represent individual characters (like your 3-digit idea), you need to use a consistent group size that covers the full range of target characters (e.g., ASCII uses 0-127, so 6 base-3 digits work since 3^6 = 729 ≥ 128). Here's how to implement this:
base = 3 group_size = 6 # Adjust this: try 5 (3^5=243) if 6 doesn't work base_data = "" with open(f"./base{base}", "r") as b3: for line in b3: base_data += line.strip("\r\n") # Pad the string with leading zeros if length isn't divisible by group size padding = group_size - (len(base_data) % group_size) if padding != group_size: base_data = "0" * padding + base_data output = [] for i in range(0, len(base_data), group_size): group = base_data[i:i+group_size] # Convert the base-3 group to a decimal value decimal_val = int(group, base) # Map to ASCII (only if value is within valid range) if 0 <= decimal_val <= 127: output.append(chr(decimal_val)) else: output.append("?") # Replace invalid characters print("".join(output))
Troubleshooting Tips
- If Solution 1 gives gibberish, try switching the
byteorderparameter frombigtolittle. - For Solution 2, test different group sizes (5, 6, 7) — check if the length of your base-3 file is evenly divisible by any of these to narrow it down.
- If decoding fails with UTF-8, try other encodings like
asciiorlatin-1.
内容的提问来源于stack exchange,提问作者Kennet Celeste

