读取文件夹中CSV文件时出现UnicodeDecodeError编码错误求助
解决Spyder中读取CSV文件时的UnicodeDecodeError问题
问题重现
你在Spyder(Python 3)里写了这段读取CSV表头的代码:
# mypath = folder directory with the csv files for each_file in listdir(mypath): with open(mypath +"//"+each_file) as f: first_line = f.readline().strip().split(",")
运行时却碰到了这个错误:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0x80 in position 3131: invalid start byte
为什么会出现这个问题?
你可能觉得自己没做任何编码操作,但实际上Python的open()函数在默认情况下,会使用你系统的默认编码去解码文件(Python 3里这个默认编码由locale.getpreferredencoding(False)返回)。而你的CSV文件并不是用UTF-8编码保存的,0x80这个字节在UTF-8编码规则里是无效的起始字节,所以解码失败了。
解决方法
这里给你几个实用的方案:
1. 指定正确的编码格式
先试试常见的非UTF-8编码,比如GBK(中文系统里很常用)、ISO-8859-1(通用兼容编码):
# 示例:用GBK编码打开 for each_file in listdir(mypath): with open(mypath +"//"+each_file, encoding='gbk') as f: first_line = f.readline().strip().split(",")
如果不确定文件编码,可以用chardet库自动检测:
- 先安装chardet:在Spyder的终端里输入
pip install chardet - 然后修改代码:
from os import listdir import chardet mypath = "你的文件夹路径" for each_file in listdir(mypath): file_path = mypath + "//" + each_file # 先读取二进制内容检测编码 with open(file_path, 'rb') as f: detect_result = chardet.detect(f.read()) # 用检测到的编码打开文件 with open(file_path, encoding=detect_result['encoding']) as f: first_line = f.readline().strip().split(",") print(f"{each_file}的表头:{first_line}")
2. 忽略或替换错误字符(不推荐,仅应急用)
如果只是临时读取表头,且不在乎少量字符丢失,可以让Python忽略错误:
with open(mypath +"//"+each_file, encoding='utf-8', errors='ignore') as f:
或者把错误字符替换成占位符�:
with open(mypath +"//"+each_file, encoding='utf-8', errors='replace') as f:
内容的提问来源于stack exchange,提问作者data_person
相关产品推荐
相关产品推荐

