You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解码三进制及其他非常规进制字符串?附Python脚本求助

Hey there! Let's break down what's going wrong with your current approach and walk through reliable solutions for converting non-standard bases (like base-3, base-5, etc.) to plaintext.

What's Wrong With Your Current Code?

Your core issue is how you're mapping base-3 digits to binary: manually converting 0→00, 1→01, 2→10 creates inconsistent value translations. For example:

  • Base-3 12 should equal 1*3 + 2 = 5 (binary 101), but your method turns it into 01 + 10 = 0110 (decimal 6). This off-by-one error cascades through the entire string, leading to garbage output.

Solution 1: Convert the Entire Base String to a Decimal Integer First

This is the most universal approach for cases where the base-encoded data represents a single large integer (similar to how base64 works under the hood). Here's the fixed code:

base = 3
base_data = ""

# Read the base-3 file content
with open(f"./base{base}", "r") as b3:
    for line in b3:
        base_data += line.strip("\r\n")

# Step 1: Convert the full base-3 string to a decimal integer
decimal_num = int(base_data, base)

# Step 2: Convert the decimal integer to a byte string
# Calculate the number of bytes needed to store the integer
byte_length = (decimal_num.bit_length() + 7) // 8
# Try 'big' first; if that fails, switch to 'little' for byte order
byte_data = decimal_num.to_bytes(byte_length, byteorder="big")

# Step 3: Decode bytes to plaintext (try different encodings if utf-8 fails)
try:
    plaintext = byte_data.decode("utf-8")
except UnicodeDecodeError:
    # Fallback to replace invalid characters, or try 'ascii'/'gbk' instead
    plaintext = byte_data.decode("utf-8", errors="replace")

print(plaintext)

Solution 2: Fixed-Length Group Mapping (If Data Is Chunked)

If the base-3 data was encoded by grouping digits to represent individual characters (like your 3-digit idea), you need to use a consistent group size that covers the full range of target characters (e.g., ASCII uses 0-127, so 6 base-3 digits work since 3^6 = 729 ≥ 128). Here's how to implement this:

base = 3
group_size = 6  # Adjust this: try 5 (3^5=243) if 6 doesn't work
base_data = ""

with open(f"./base{base}", "r") as b3:
    for line in b3:
        base_data += line.strip("\r\n")

# Pad the string with leading zeros if length isn't divisible by group size
padding = group_size - (len(base_data) % group_size)
if padding != group_size:
    base_data = "0" * padding + base_data

output = []
for i in range(0, len(base_data), group_size):
    group = base_data[i:i+group_size]
    # Convert the base-3 group to a decimal value
    decimal_val = int(group, base)
    # Map to ASCII (only if value is within valid range)
    if 0 <= decimal_val <= 127:
        output.append(chr(decimal_val))
    else:
        output.append("?")  # Replace invalid characters

print("".join(output))

Troubleshooting Tips

  • If Solution 1 gives gibberish, try switching the byteorder parameter from big to little.
  • For Solution 2, test different group sizes (5, 6, 7) — check if the length of your base-3 file is evenly divisible by any of these to narrow it down.
  • If decoding fails with UTF-8, try other encodings like ascii or latin-1.

内容的提问来源于stack exchange,提问作者Kennet Celeste

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:49:10