You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CSV文件触发UnicodeDecodeError报错,请求排查问题并运行程序

UnicodeDecodeError: 'utf-8' codec解码失败问题解析

错误信息

UnicodeDecodeError: 'utf-8' codec can't decode bytes in position 606-607: invalid continuation byte

错误原因

该错误发生在尝试解码不符合UTF-8规范的字节序列时,常见场景包括:

  • 读取的文件采用了GBK、GB2312等非UTF-8编码格式
  • 接收的网络数据或二进制文件中包含不符合UTF-8规则的字节片段

触发错误的示例程序

假设存在一个用GBK编码的文件non_utf8_file.txt,执行以下代码会触发报错:

with open("non_utf8_file.txt", "r", encoding="utf-8") as f:
    content = f.read()

解决方法

  1. 使用正确编码解码
    确认文件/数据的实际编码,用对应编码读取:
# 示例:文件实际编码为GBK时
with open("non_utf8_file.txt", "r", encoding="gbk") as f:
    content = f.read()
  1. 添加错误处理策略
    如果无法确定编码,可以选择忽略或替换无效字节:
# 忽略无效字节
with open("non_utf8_file.txt", "r", encoding="utf-8", errors="ignore") as f:
    content = f.read()

# 用?替换无效字节
with open("non_utf8_file.txt", "r", encoding="utf-8", errors="replace") as f:
    content = f.read()

内容的提问来源于stack exchange,提问作者Subhankar Rout

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 17:02:38