You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3.7.3解码HDT格式中CESU-8编码字节串问题求助

Why can't Python decode this DBpedia byte string, but a web tool can?

I ran into a tricky encoding issue with a byte string from DBpedia's HDT compressed format, and figured out the root cause and fix—sharing this to help others who hit the same problem.

The Problem

I had this byte string:

b'"\xc2\xb7\xed\xa0\x81\xed\xb1\x96\xed\xa0\x81\xed\xb1\xb1\xed\xa0\x81\xed\xb1\x9d\xed\xa0\x81\xed\xb1\xbe\xed\xa0\x81\xed\xb1\xaf \xed\xa0\x81\xed\xb1\xa9\xed\xa0\x81\xed\xb1\xa4\xed\xa0\x81\xed\xb1\x93\xed\xa0\x81\xed\xb1\xa9\xed\xa0\x81\xed\xb1\x9a\xed\xa0\x81\xed\xb1\xa7\xed\xa0\x81\xed\xb1\x91"@en'

When I used an online UTF-8 decoder, it correctly decoded to:
"·іѱѝѾѯ ѩѤѓѩњѧё"@en

But in Python 3.7.3, calling mystring.decode('utf8') threw a UnicodeDecodeError:

UnicodeDecodeError: 'utf-8' codec can't decode byte 0xed in position 3: invalid continuation byte

To make things weirder, encoding the target string back to UTF-8 gave a completely different byte string that decoded fine in Python.

The Root Cause

The original byte string uses CESU-8, a deprecated UTF-8 variant. Here's how it differs from standard UTF-8:

  • Standard UTF-8 encodes Unicode characters outside the Basic Multilingual Plane (BMP) with a single 4-byte sequence.
  • CESU-8 splits these characters into UTF-16 surrogate pairs first, then encodes each surrogate as a 3-byte UTF-8 sequence. This creates a non-standard byte stream that Python's strict UTF-8 decoder rejects.

Online decoding tools often include logic to handle these non-standard variants for better compatibility, which is why they worked where Python didn't.

The Fix

Use the python-ftfy library, which has built-in support for decoding UTF-8 variants like CESU-8. Specifically, use the utf-8-variants encoding:

import ftfy

mystring = b'"\xc2\xb7\xed\xa0\x81\xed\xb1\x96\xed\xa0\x81\xed\xb1\xb1\xed\xa0\x81\xed\xb1\x9d\xed\xa0\x81\xed\xb1\xbe\xed\xa0\x81\xed\xb1\xaf \xed\xa0\x81\xed\xb1\xa9\xed\xa0\x81\xed\xb1\xa4\xed\xa0\x81\xed\xb1\x93\xed\xa0\x81\xed\xb1\xa9\xed\xa0\x81\xed\xb1\x9a\xed\xa0\x81\xed\xb1\xa7\xed\xa0\x81\xed\xb1\x91"@en'
decoded_string = ftfy.decode(mystring, encoding='utf-8-variants')
print(decoded_string)  # Output: "·іѱѝѾѯ ѩѤѓѩњѧё"@en

This will correctly parse the CESU-8 encoded byte string into the intended Unicode text.

内容的提问来源于stack exchange,提问作者Folkvir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:19:48