读取含特殊编码字符的CSV时cp1251转utf-8出现UnicodeEncodeError如何解决
解决方案
问题根因
调用codecs.open读取文件时指定了encoding='cp1251',读取过程中会自动将文件原始字节用cp1251解码为Unicode字符串,若原始文件存在cp1251不支持的字节,会被自动替换为占位符\ufffd。该占位符不存在于cp1251编码表中,后续对字符串执行encode('cp1251')操作时就会触发编码错误。
转换逻辑应直接处理文件原始字节,无需提前执行解码操作。
修正代码
with open("dataset_classif_chatbot.csv", "rb") as in_f, open("dataset_classif_chatbot_cleaned.csv", "w", encoding="utf-8") as out_f: for line_bytes in in_f: # 按cp1251解码得到乱码字符串,再编码回原始utf-8字节后解码为正确文本 corrected_line = line_bytes.decode("cp1251").encode("cp1251").decode("utf-8") out_f.write(corrected_line)
异常处理
如果运行时仍有报错,可给编码步骤添加错误处理参数,避免程序中断:
# 遇到不支持的字符时用?替换 corrected_line = line_bytes.decode("cp1251").encode("cp1251", errors="replace").decode("utf-8") # 遇到不支持的字符时直接忽略 corrected_line = line_bytes.decode("cp1251").encode("cp1251", errors="ignore").decode("utf-8")
内容的提问来源于stack exchange,提问作者Notsogenius
相关产品推荐
相关产品推荐

