Python3.7.3解码HDT格式中CESU-8编码字节串问题求助
I ran into a tricky encoding issue with a byte string from DBpedia's HDT compressed format, and figured out the root cause and fix—sharing this to help others who hit the same problem.
The Problem
I had this byte string:
b'"\xc2\xb7\xed\xa0\x81\xed\xb1\x96\xed\xa0\x81\xed\xb1\xb1\xed\xa0\x81\xed\xb1\x9d\xed\xa0\x81\xed\xb1\xbe\xed\xa0\x81\xed\xb1\xaf \xed\xa0\x81\xed\xb1\xa9\xed\xa0\x81\xed\xb1\xa4\xed\xa0\x81\xed\xb1\x93\xed\xa0\x81\xed\xb1\xa9\xed\xa0\x81\xed\xb1\x9a\xed\xa0\x81\xed\xb1\xa7\xed\xa0\x81\xed\xb1\x91"@en'
When I used an online UTF-8 decoder, it correctly decoded to:"·іѱѝѾѯ ѩѤѓѩњѧё"@en
But in Python 3.7.3, calling mystring.decode('utf8') threw a UnicodeDecodeError:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xed in position 3: invalid continuation byte
To make things weirder, encoding the target string back to UTF-8 gave a completely different byte string that decoded fine in Python.
The Root Cause
The original byte string uses CESU-8, a deprecated UTF-8 variant. Here's how it differs from standard UTF-8:
- Standard UTF-8 encodes Unicode characters outside the Basic Multilingual Plane (BMP) with a single 4-byte sequence.
- CESU-8 splits these characters into UTF-16 surrogate pairs first, then encodes each surrogate as a 3-byte UTF-8 sequence. This creates a non-standard byte stream that Python's strict UTF-8 decoder rejects.
Online decoding tools often include logic to handle these non-standard variants for better compatibility, which is why they worked where Python didn't.
The Fix
Use the python-ftfy library, which has built-in support for decoding UTF-8 variants like CESU-8. Specifically, use the utf-8-variants encoding:
import ftfy mystring = b'"\xc2\xb7\xed\xa0\x81\xed\xb1\x96\xed\xa0\x81\xed\xb1\xb1\xed\xa0\x81\xed\xb1\x9d\xed\xa0\x81\xed\xb1\xbe\xed\xa0\x81\xed\xb1\xaf \xed\xa0\x81\xed\xb1\xa9\xed\xa0\x81\xed\xb1\xa4\xed\xa0\x81\xed\xb1\x93\xed\xa0\x81\xed\xb1\xa9\xed\xa0\x81\xed\xb1\x9a\xed\xa0\x81\xed\xb1\xa7\xed\xa0\x81\xed\xb1\x91"@en' decoded_string = ftfy.decode(mystring, encoding='utf-8-variants') print(decoded_string) # Output: "·іѱѝѾѯ ѩѤѓѩњѧё"@en
This will correctly parse the CESU-8 encoded byte string into the intended Unicode text.
内容的提问来源于stack exchange,提问作者Folkvir

