You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取含特殊编码字符的CSV时cp1251转utf-8出现UnicodeEncodeError如何解决

解决方案

问题根因

调用codecs.open读取文件时指定了encoding='cp1251',读取过程中会自动将文件原始字节用cp1251解码为Unicode字符串,若原始文件存在cp1251不支持的字节,会被自动替换为占位符\ufffd。该占位符不存在于cp1251编码表中,后续对字符串执行encode('cp1251')操作时就会触发编码错误。
转换逻辑应直接处理文件原始字节,无需提前执行解码操作。

修正代码

with open("dataset_classif_chatbot.csv", "rb") as in_f, open("dataset_classif_chatbot_cleaned.csv", "w", encoding="utf-8") as out_f:
    for line_bytes in in_f:
        # 按cp1251解码得到乱码字符串,再编码回原始utf-8字节后解码为正确文本
        corrected_line = line_bytes.decode("cp1251").encode("cp1251").decode("utf-8")
        out_f.write(corrected_line)

异常处理

如果运行时仍有报错,可给编码步骤添加错误处理参数,避免程序中断:

# 遇到不支持的字符时用?替换
corrected_line = line_bytes.decode("cp1251").encode("cp1251", errors="replace").decode("utf-8")
# 遇到不支持的字符时直接忽略
corrected_line = line_bytes.decode("cp1251").encode("cp1251", errors="ignore").decode("utf-8")

内容的提问来源于stack exchange,提问作者Notsogenius

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 02:06:07