在Colab中读取CSV文件时遭遇Python错误求助
问题描述
已通过以下代码在Google Colab中挂载Google Drive:
from google.colab import drive drive.mount('/content/gdrive')
尝试用下方代码读取存储在Drive中的weather.csv文件:
weather = pd.read_csv('gdrive/My Drive/weather.csv', sep=',', encoding='latin-1', header=None)
但触发ParserError:
ParserError Traceback (most recent call last) <ipython-input-101-d66270414100> in <cell line: 1>() ----> 1 weather=pd.read_csv('gdrive/My Drive/weather.csv', sep=',', encoding='latin-1', header=None) 9 frames /usr/local/lib/python3.10/dist-packages/pandas/_libs/parsers.pyx in pandas._libs.parsers.raise_parser_error() ParserError: Error tokenizing data. C error: Expected 1 fields in line 4, saw 3
调整分隔符、编码、header参数后仍报错;不指定编码时会触发UnicodeDecodeError:
UnicodeDecodeError Traceback (most recent call last) <ipython-input-102-258bcb4bc6f5> in <cell line: 1>() ----> 1 weather=pd.read_csv('gdrive/My Drive/weather.csv', sep=',', header=None) 10 frames /usr/local/lib/python3.10/dist-packages/pandas/_libs/parsers.pyx in pandas._libs.parsers.raise_parser_error() UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe3 in position 14: invalid continuation byte
文件前几行内容如下:
Wed 17 Nov 2021 03:20:03 PM GMT, 24.2, 9.8, 35.780, 1, 26.11, 1006.79, 31.14 Wed 17 Nov 2021 03:21:02 PM GMT, 24.2, 9.8, 36.318, 1, 26.22, 1006.76, 31.03 Wed 17 Nov 2021 03:22:03 PM GMT, 24.2, 9.8, 36.856, 1, 26.24, 1006.80, 30.98 Wed 17 Nov 2021 03:23:05 PM GMT, 24.2, 9.8, 40.84, 1, 26.28, 1006.84, 30.91 Wed 17 Nov 2021 03:24:02 PM GMT, 24.2, 9.8, 36.856, 1, 26.47, 1006.86, 30.77 Wed 17 Nov 2021 03:25:03 PM GMT, 24.2, 9.8, 36.856, 1, 26.59, 1006.90, 30.57 Wed 17 Nov 2021 03:26:02 PM GMT, 24.2, 9.6, 36.318, 1, 26.63, 1006.92, 30.51 Wed 17 Nov 2021 03:27:03 PM GMT, 24.2, 9.6, 36.856, 1, 26.69, 1006.92, 30.54
解决方案
结合报错信息和文件内容,问题源于文件编码不匹配和部分行格式异常的组合,可按以下步骤处理:
1. 检测文件真实编码
先用chardet工具检测文件的实际编码(Colab需先安装):
!pip install chardet import chardet # 读取文件前1000字节检测编码 with open('/content/gdrive/My Drive/weather.csv', 'rb') as f: encoding_info = chardet.detect(f.read(1000)) print(encoding_info) # 输出示例:{'encoding': 'ISO-8859-1', 'confidence': 0.73, 'language': ''}
根据输出的encoding值指定读取时的编码参数,比如检测到ISO-8859-1可直接用encoding='latin-1'(两者等价)。
2. 跳过格式异常的行
ParserError提示第4行字段数不符,说明文件中存在换行错误或分隔符异常的行,添加on_bad_lines='skip'参数跳过这些行:
import pandas as pd # 替换为检测到的编码,比如'latin-1' weather = pd.read_csv('/content/gdrive/My Drive/weather.csv', sep=',', encoding='latin-1', header=None, on_bad_lines='skip')
若不想直接跳过,可改用on_bad_lines='warn'查看异常行的具体内容,再针对性修复文件。
3. 使用绝对路径读取
避免简写路径可能带来的问题,直接使用Colab挂载后的绝对路径/content/gdrive/My Drive/weather.csv。
4. 强制指定列数(可选)
从提供的文件内容看每行有8列,可指定usecols=range(8)强制读取前8列,避免列数异常干扰:
weather = pd.read_csv('/content/gdrive/My Drive/weather.csv', sep=',', encoding='latin-1', header=None, usecols=range(8), on_bad_lines='skip')
内容的提问来源于stack exchange,提问作者Louisa Wareing
相关产品推荐
相关产品推荐

