Pandas DataFrame触发KeyError: 'date'问题,求解决方案
问题描述
已查阅两个Stack Overflow上关于KeyError: 'date'的相关问题,但未能解决问题。执行代码将date列转为datetime格式时触发KeyError: 'date',无额外解释。
代码
import pandas as pd, numpy as np import csv import warnings from bs4 import BeautifulSoup, MarkupResemblesLocatorWarning from sklearn.impute import SimpleImputer from sklearn.exceptions import ConvergenceWarning from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.preprocessing import LabelEncoder from sklearn.linear_model import LinearRegression, LogisticRegression, Perceptron from sklearn.tree import DecisionTreeClassifier from sklearn.metrics import mean_squared_error, r2_score, accuracy_score, confusion_matrix, ConfusionMatrixDisplay import seaborn as sns import matplotlib.pyplot as plt ## 读取数据 dtypes = { 'Unnamed: 0': 'int32', 'drugName': 'category', 'condition': 'category', 'review': 'category', 'rating': 'float16', 'date': 'categorical', 'usefulCount': 'int16' } train_df = pd.read_csv('/content/drugsComTrain_raw.tsv', sep='\t', quoting=2, dtype=dtypes) # 随机选取训练集的80%数据 train_df = train_df.sample(frac=0.8, random_state=42) test_df = pd.read_csv('/content/drugsComTest_raw.tsv', sep='\t', quoting=2, dtype=dtypes) print(train_df.head()) ## 将date列转为datetime格式 train_df['date'], test_df['date'] = pd.to_datetime(train_df['date'], format='%b %d, %Y'), pd.to_datetime(test_df['date'], format='%b %d, %Y') # 此行触发报错
报错信息
KeyError Traceback (most recent call last) /usr/local/lib/python3.10/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3801 try: -> 3802 return self._engine.get_loc(casted_key) 3803 except KeyError as err: 4 frames /usr/local/lib/python3.10/dist-packages/pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc() /usr/local/lib/python3.10/dist-packages/pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc() pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item() pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item() KeyError: 'date' The above exception was the direct cause of the following exception: KeyError Traceback (most recent call last) <ipython-input-17-056c9fab2e6c> in <cell line: 24>() 22 print(train_df.head()) 23 ## 将date列转为datetime格式 ---> 24 train_df['date'], test_df['date'] = pd.to_datetime(train_df['date'], format='%b %d, %Y'), pd.to_datetime(test_df['date'], format='%b %d, %Y') 25 26 ## 提取日、月、年到单独列 /usr/local/lib/python3.10/dist-packages/pandas/core/frame.py in __getitem__(self, key) 3805 if self.columns.nlevels > 1: 3806 return self._getitem_multilevel(key) -> 3807 indexer = self.columns.get_loc(key) 3808 if is_integer(indexer): 3809 indexer = [indexer] /usr/local/lib/python3.10/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3802 return self._engine.get_loc(casted_key) 3803 except KeyError as err: -> 3804 raise KeyError(key) from err 3805 except TypeError: 3806 # 如果是类列表的键,_check_indexing_error会抛出异常 KeyError: 'date'
解决方案
1. 排查列名问题
先执行以下代码,确认DataFrame的实际列名是否存在拼写、大小写、引号或空格异常:
print("训练集列名:", train_df.columns.tolist()) print("测试集列名:", test_df.columns.tolist())
2. 清理异常列名
如果输出的列名带有双引号(比如'"date"')或首尾空格,执行以下代码清理:
# 去除列名的双引号和首尾空格 train_df.columns = train_df.columns.str.strip('"').str.strip() test_df.columns = test_df.columns.str.strip('"').str.strip()
这是因为你读文件时设置了quoting=2(即csv.QUOTE_ALL),会导致列名被双引号包裹,从而无法通过'date'正常访问。
3. 重新转换日期格式
列名清理完成后,再执行日期转换代码:
train_df['date'] = pd.to_datetime(train_df['date'], format='%b %d, %Y') test_df['date'] = pd.to_datetime(test_df['date'], format='%b %d, %Y')
可选优化:调整读文件参数
后续读取文件时,可去掉quoting=2(或设置为quoting=0即csv.QUOTE_MINIMAL),避免列名被不必要的引号包裹:
train_df = pd.read_csv('/content/drugsComTrain_raw.tsv', sep='\t', quoting=0, dtype=dtypes) test_df = pd.read_csv('/content/drugsComTest_raw.tsv', sep='\t', quoting=0, dtype=dtypes)
内容的提问来源于stack exchange,提问作者Karthik Bhandary
相关产品推荐
相关产品推荐

