You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决Python读取CSV时'utf-8'解码0x99字节失败问题求助

解决Pandas读取CSV时的UTF-8解码错误问题

你遇到的'utf-8' codec can't decode byte 0x99错误,其实是因为文件里包含了UTF-8不支持的特殊字符——0x99这个字节在Windows-1252编码里代表商标符号™,这是很多Windows环境下导出的CSV文件常用的编码,而非你尝试的ISO-8859-1或UTF-8。下面给你几个可行的解决方案:

方案1:直接指定Windows-1252编码读取

这是最直接的尝试,因为0x99是Windows-1252的典型特殊字符:

import pandas as pd

# 替换为你的文件路径
file_path = 'D:\\DJ\\Placement reports\\Copy of Placement Reports _ Apr_Mar_May Page 2.csv'
rawdata = pd.read_csv(file_path, encoding='windows-1252')

方案2:用chardet自动检测文件编码

如果不确定编码,用chardet库自动检测是最稳妥的方法:

  1. 先安装chardet(如果没装的话):
pip install chardet
  1. 编写检测和读取代码:
import chardet
import pandas as pd

file_path = 'D:\\DJ\\Placement reports\\Copy of Placement Reports _ Apr_Mar_May Page 2.csv'

# 读取文件前10KB内容用于编码检测(足够覆盖特征字符)
with open(file_path, 'rb') as f:
    detect_result = chardet.detect(f.read(10000))

print(f"检测到的编码: {detect_result['encoding']}")
print(f"检测置信度: {detect_result['confidence']}")

# 使用检测出的编码读取文件
rawdata = pd.read_csv(file_path, encoding=detect_result['encoding'])

方案3:临时忽略错误(不推荐,仅用于排查)

如果只是想先加载文件排查问题,可以用errors='ignore'参数跳过解码错误,但会丢失特殊字符,不建议用于正式分析:

rawdata = pd.read_csv(file_path, encoding='utf-8', errors='ignore')

补充说明

你之前尝试的sys.setdefaultencoding("ISO-8859-1")在Python 2.x里可能生效,但ISO-8859-1会把0x99解析成乱码(比如™),而Windows-1252是ISO-8859-1的超集,能正确解析这类Windows环境下的特殊字符。如果你的环境是Python 3.x,这个设置本身就无效,无需使用。

内容的提问来源于stack exchange,提问作者Debjyoti Das

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:00:28