使用Pandas导入500MB CSV文件时遇UnicodeDecodeError错误求助
解决Pandas导入CSV时的UnicodeDecodeError错误
问题场景
尝试用Pandas导入500MB的CSV文件,执行代码:
import pandas as pd df = pd.read_csv('filename.csv') df.head()
运行后报错:
Traceback (most recent call last): File "/Users/Filename.py", line 3, in <module> df = pd.read_csv ('/Users/Filename.csv') File "/Users/venv/lib/python3.9/site-packages/pandas/util/_decorators.py", line 211, in wrapper return func(*args, **kwargs) File "/Users/venv/lib/python3.9/site-packages/pandas/util/_decorators.py", line 331, in wrapper return func(*args, **kwargs) File "/Users/venv/lib/python3.9/site-packages/pandas/io/parsers/readers.py", line 950, in read_csv return _read(filepath_or_buffer, kwds) File "/Usersvenv/lib/python3.9/site-packages/pandas/io/parsers/readers.py", line 605, in _read parser = TextFileReader(filepath_or_buffer, **kwds) File "/Users/venv/lib/python3.9/site-packages/pandas/io/parsers/readers.py", line 1442, in __init__ self._engine = self._make_engine(f, self.engine) File "/Users/venv/lib/python3.9/site-packages/pandas/io/parsers/readers.py", line 1753, in _make_engine return mapping[engine](f, **self.options) File "/Users/venv/lib/python3.9/site-packages/pandas/io/parsers/c_parser_wrapper.py", line 79, in __init__ self._reader = parsers.TextReader(src, **kwds) File "pandas/_libs/parsers.pyx", line 547, in pandas._libs.parsers.TextReader.__cinit__ File "pandas/_libs/parsers.pyx", line 636, in pandas._libs.parsers.TextReader._get_header File "pandas/_libs/parsers.pyx", line 852, in pandas._libs.parsers.TextReader._tokenize_rows File "pandas/_libs/parsers.pyx", line 1965, in pandas._libs.parsers.raise_parser_error UnicodeDecodeError: 'utf-8' codec can't decode byte 0xa5 in position 4540: invalid start byte
原因
该错误说明你的CSV文件并非UTF-8编码,而Pandas默认使用UTF-8解码,因此遇到无法识别的字节时触发报错。0xa5字节常见于GBK/GB2312等中文编码(对应字符为“¥”)。
解决方法
方法1:指定正确编码格式
直接在read_csv中传入文件对应的编码参数,比如GBK:
import pandas as pd df = pd.read_csv('filename.csv', encoding='gbk') df.head()
如果GBK无效,可尝试gb2312、cp1252、utf-16等其他常见编码。
方法2:自动检测文件编码
借助chardet库自动识别文件编码:
- 先安装依赖库:
pip install chardet
- 检测编码并读取文件:
import pandas as pd import chardet # 读取文件前100KB内容用于编码检测 with open('filename.csv', 'rb') as f: detect_result = chardet.detect(f.read(100000)) # 使用检测到的编码读取文件 df = pd.read_csv('filename.csv', encoding=detect_result['encoding']) df.head()
方法3:忽略解码错误(不推荐)
若无需保留无法解码的字符,可添加errors='ignore'参数跳过错误,但会丢失部分数据:
df = pd.read_csv('filename.csv', encoding='utf-8', errors='ignore')
方法4:分块读取大文件
针对500MB的大文件,即使编码正确也可能占用过多内存,建议分块读取:
import pandas as pd chunk_size = 10000 # 每次读取10000行 chunk_list = [] # 逐块读取并存储 for chunk in pd.read_csv('filename.csv', encoding='gbk', chunksize=chunk_size): chunk_list.append(chunk) # 合并所有块为完整DataFrame df = pd.concat(chunk_list, ignore_index=True) df.head()
内容的提问来源于stack exchange,提问作者colotech322
相关产品推荐
相关产品推荐

