Python2.7中Unicode编码与urllib.urlencode报错的技术问询
Hey there, let's unpack your Python 2.7 Unicode confusion—this stuff was a major pain point back in the day, so I totally get why you're digging into the underlying logic. Let's tackle each of your questions one by one:
1. What does str() do when called on a UTF-8 encoded value?
In Python 2, when you encode a Unicode object with .encode('utf-8'), you get a str type object (which is a sequence of bytes, not abstract characters). When you call str() on this already-encoded str object, it does basically nothing—it just returns the same byte sequence.
Compare this to calling str() on a Unicode object: that triggers an implicit encoding step, where Python tries to convert the abstract Unicode characters into bytes using the system's default encoding (which in your case is ASCII). Since your Unicode string has characters outside the ASCII range, that's why you hit the UnicodeEncodeError.
Example to clarify:
# Unicode object (abstract characters) unicode_str = u'some_v盲l眉e' # UTF-8 encoded byte string (str type in Python 2) utf8_bytes = unicode_str.encode('utf-8') type(utf8_bytes) # Returns <type 'str'> str(utf8_bytes) == utf8_bytes # Returns True—no change happens
2. Does the representation change when a Unicode object is encoded to UTF-8?
Absolutely—they're fundamentally different types with distinct representations:
- A Unicode object (
u''syntax in Python 2) stores abstract characters. Think of it as a list of "character IDs" (code points) that represent what the text is, not how it's stored as bytes. - A UTF-8 encoded
strobject stores raw bytes. Each character in the original Unicode string is converted to one or more bytes according to UTF-8 rules. For non-ASCII characters (like your盲or眉), this means multiple bytes instead of a single ASCII byte.
For example, if you print both:
print(u'some_v盲l眉e') # Prints the actual characters (if your terminal supports it) print(u'some_v盲l眉e'.encode('utf-8')) # Prints the byte sequence, which might look like b'some_v\xe7\x9b\xb2l\xe7\x9c\x89e' (depending on your setup)
The underlying data is completely different: one is a sequence of code points, the other is a sequence of bytes.
3. Why does str() try to encode Unicode objects to ASCII by default, if the representation changes?
This is a historical design choice in Python 2. When Python 2 was created, ASCII was the dominant encoding, so the language defaulted to using it for implicit conversions between Unicode and byte strings.
The str() function, when given a Unicode object, is equivalent to calling unicode_obj.encode(sys.getdefaultencoding()). Since your system's default encoding is ASCII, Python tries to convert every character in the Unicode string to an ASCII byte—but characters like 盲 or 眉 don't have ASCII equivalents, hence the error.
This is one of the biggest pitfalls of Python 2's Unicode handling, which is why Python 3 completely separated str (Unicode characters) and bytes (raw bytes) to avoid these implicit conversion surprises.
内容的提问来源于stack exchange,提问作者Kirk Douglas Jr

