Python与Node.js的Base64解码差异问题求助
Hey there, let’s figure out why you’re seeing mismatched results between Python and Node.js when decoding that Base64 string. The good news is this isn’t a bug in either language—it’s just a difference in how they handle string encoding out of the box.
Let’s Start with the Facts
First, let’s recap the code and outputs you shared to keep everyone on the same page:
Node.js Code:
console.log(Buffer.from('Im3Osc6_z4HPgc-J==', 'base64').toString());Output:
"mαορρω"Python Code (using the deprecated
decodestringyou provided):from base64 import decodestring print decodestring('Im3Osc6_z4HPgc-J==')Output:
"mαγ?s?p"
The Root Cause: Encoding Defaults
Both languages decode the Base64 string to the exact same raw byte sequence—let’s confirm that first:
The Base64 string Im3Osc6_z4HPgc-J== decodes to these hex bytes: 22 6D CE B1 CE AC CF 81 CF 89
The difference comes in how each language turns those bytes into a human-readable string:
Node.js: When you call
.toString()on a Buffer without specifying an encoding, it automatically uses UTF-8. Let’s map those bytes to UTF-8 characters:6D→'m'CE B1→ Greek letter α (U+03B1)CE AC→ Greek letter ο (U+03AC)CF 81→ Greek letter ρ (U+03C1)CF 89→ Greek letter ω (U+03C9)
That’s exactly the clean output you see from Node.js.
Python: The
decodestringfunction returns a byte string (in Python 2, this is astr; in Python 3, it’s abytesobject). When you print this directly, Python doesn’t assume UTF-8 by default:- In Python 2, it uses your terminal’s default encoding (which might be something like Latin-1 instead of UTF-8). Multi-byte UTF-8 sequences get split into single-byte characters that don’t make sense, or invalid characters get replaced with
?. - In Python 3, printing a
bytesobject will show you the raw byte representation, not the decoded UTF-8 string.
- In Python 2, it uses your terminal’s default encoding (which might be something like Latin-1 instead of UTF-8). Multi-byte UTF-8 sequences get split into single-byte characters that don’t make sense, or invalid characters get replaced with
The Fix for Python
To get the same correct output as Node.js, you need to explicitly decode the bytes to a UTF-8 string. Here’s how to do it for both Python versions:
Python 2
from base64 import decodestring decoded_bytes = decodestring('Im3Osc6_z4HPgc-J==') # Explicitly decode as UTF-8 print(decoded_bytes.decode('utf-8'))
Python 3 (note: decodestring is deprecated—use b64decode instead)
from base64 import b64decode decoded_bytes = b64decode('Im3Osc6_z4HPgc-J==') # Decode bytes to UTF-8 string print(decoded_bytes.decode('utf-8'))
Either of these will give you the expected "mαορρω" output.
Quick Recap
Node.js defaults to UTF-8 when converting Buffer data to a string, but Python requires you to be explicit about decoding bytes to UTF-8. Once you add that explicit decode step, the results match perfectly.
内容的提问来源于stack exchange,提问作者Qiang Li

