You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理EUCKR转UTF-8乱码:将无法转换字符替换为?

解决EUCKR转UTF-8时的解码错误问题

问题根源是解码时默认使用strict模式,遇到非法字节序列直接抛出异常。只需在decode方法中添加errors='replace'参数,就能将无法解码的字符自动替换为?,避免报错。

修改关键代码行

将原代码中的解码逻辑:

el.encode('latin1').decode('euc-kr')

修改为:

el.encode('latin1').decode('euc-kr', errors='replace')

修改后的完整代码

cursor = conn.cursor()
cursor.execute("select * from %s" % table_name)

row = cursor.fetchall()
# 加入errors='replace'处理无法解码的字符
data = [tuple(el.encode('latin1').decode('euc-kr', errors='replace') for el in t) for t in row]

# Open CSV file for writing.
csvFile = csv.writer(open(filePath + fileName, 'w', newline='', encoding='utf-8'),
                    delimiter=',', lineterminator='\r\n',
                    quoting=csv.QUOTE_ALL, escapechar='\\')

csvFile.writerows(data)

补充说明

  • errors='replace'是Python字符串解码的标准参数,除了replace还有ignore(忽略错误字符)、strict(默认,抛出异常)等选项,可根据需求选择。
  • 需确保从数据库获取的字符串el确实是以Latin1编码存储的,这样先编码为Latin1字节流、再用EUCKR解码的流程才是正确的。

内容的提问来源于stack exchange,提问作者Bellpump

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 09:22:24