Python 2.7与JS中encode/decode含义差异及中文转UTF-8困惑
encode/decode Behave Differently in JavaScript vs. Python 2.7 Great question! The confusion here boils down to a fundamental difference in how each language represents strings under the hood. Let's break this down step by step:
Core String Model Differences
First, let's clarify the string types each language uses—this is the key to understanding the method names:
- JavaScript: All strings are inherently Unicode character sequences. There's no separate "byte string" type; every string you work with is a collection of Unicode code points.
- Python 2.7: There are two distinct string types:
str: A sequence of raw bytes (already encoded in some character set like UTF-8, ASCII, or GBK).unicode: A sequence of Unicode code points (the "raw" character representation, unencoded).
What encode/decode Mean in Each Language
JavaScript
In JS, since all strings are Unicode, the encodeXxx methods (like encodeURIComponent) do something specific:
- They take the Unicode string and convert it into a URI-safe escaped string (e.g., spaces become
%20, non-ASCII characters become%XXsequences). This is effectively converting Unicode characters to a byte-based representation that's safe for URIs, wrapped in a string. - Conversely,
decodeXxxmethods take those escaped strings and convert them back to regular Unicode strings.
Python 2.7
Python 2's encode and decode map directly to the conversion between its two string types:
decode(): Converts a byte string (str) to a Unicode string (unicode). This is "decoding" bytes into the characters they represent, using the specified encoding (e.g., UTF-8). So when you run"奥多比".decode("utf-8"), you're telling Python: "Take these UTF-8 bytes and turn them into the corresponding Unicode characters"—which gives youu'\u5965\u591a\u6bd4'.encode(): Converts a Unicode string (unicode) to a byte string (str). This is "encoding" characters into a specific byte format (like UTF-8). So you'd use this on aunicodestring, e.g.,u"奥多比".encode("utf-8")gives you the UTF-8 byte sequence'\xe5\xa5\xa5\xe5\xa4\x9a\xe6\xaf\x94'.
Why encode("utf-8") Doesn't Work on "奥多比" in Python 2.7
The string "奥多比" in Python 2 is already a str type (a byte string). If you call encode("utf-8") on it, Python does something unexpected:
- It first tries to implicitly decode the byte string to
unicodeusing your system's default encoding (which might not be UTF-8!). - Then it encodes that
unicodestring to UTF-8 bytes.
This is not only unnecessary but often causes errors (e.g., if your system encoding doesn't match the actual encoding of the byte string). To get the result you want, you need to start with a unicode string if you're going to use encode().
Quick Cheat Sheet
| Operation | JavaScript | Python 2.7 |
|---|---|---|
| Unicode chars → byte representation | encodeURIComponent(str) | unicode_str.encode("utf-8") |
| Byte representation → Unicode chars | decodeURIComponent(str) | byte_str.decode("utf-8") |
内容的提问来源于stack exchange,提问作者Bhumi Singhal

