You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame触发KeyError: 'date'问题,求解决方案

问题描述

已查阅两个Stack Overflow上关于KeyError: 'date'的相关问题,但未能解决问题。执行代码将date列转为datetime格式时触发KeyError: 'date',无额外解释。

代码
import pandas as pd, numpy as np
import csv
import warnings
from bs4 import BeautifulSoup, MarkupResemblesLocatorWarning
from sklearn.impute import SimpleImputer
from sklearn.exceptions import ConvergenceWarning
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.preprocessing import LabelEncoder
from sklearn.linear_model import LinearRegression, LogisticRegression, Perceptron
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import mean_squared_error, r2_score, accuracy_score, confusion_matrix, ConfusionMatrixDisplay 
import seaborn as sns
import matplotlib.pyplot as plt

## 读取数据
dtypes = { 'Unnamed: 0': 'int32', 'drugName': 'category', 'condition': 'category', 'review': 'category', 'rating': 'float16', 'date': 'categorical', 'usefulCount': 'int16' }
train_df = pd.read_csv('/content/drugsComTrain_raw.tsv', sep='\t', quoting=2, dtype=dtypes)
# 随机选取训练集的80%数据
train_df = train_df.sample(frac=0.8, random_state=42)
test_df = pd.read_csv('/content/drugsComTest_raw.tsv', sep='\t', quoting=2, dtype=dtypes)

print(train_df.head())
## 将date列转为datetime格式
train_df['date'], test_df['date'] = pd.to_datetime(train_df['date'], format='%b %d, %Y'), pd.to_datetime(test_df['date'], format='%b %d, %Y') # 此行触发报错
报错信息
KeyError                                  Traceback (most recent call last)
/usr/local/lib/python3.10/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
   3801             try:
-> 3802                 return self._engine.get_loc(casted_key)
   3803             except KeyError as err:

4 frames
/usr/local/lib/python3.10/dist-packages/pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()

/usr/local/lib/python3.10/dist-packages/pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'date'

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
<ipython-input-17-056c9fab2e6c> in <cell line: 24>()
     22 print(train_df.head())
     23 ## 将date列转为datetime格式
---> 24 train_df['date'], test_df['date'] = pd.to_datetime(train_df['date'], format='%b %d, %Y'), pd.to_datetime(test_df['date'], format='%b %d, %Y')
     25 
     26 ## 提取日、月、年到单独列

/usr/local/lib/python3.10/dist-packages/pandas/core/frame.py in __getitem__(self, key)
   3805             if self.columns.nlevels > 1:
   3806                 return self._getitem_multilevel(key)
-> 3807             indexer = self.columns.get_loc(key)
   3808             if is_integer(indexer):
   3809                 indexer = [indexer]

/usr/local/lib/python3.10/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
   3802                 return self._engine.get_loc(casted_key)
   3803             except KeyError as err:
-> 3804                 raise KeyError(key) from err
   3805             except TypeError:
   3806                 # 如果是类列表的键,_check_indexing_error会抛出异常

KeyError: 'date'
解决方案

1. 排查列名问题

先执行以下代码,确认DataFrame的实际列名是否存在拼写、大小写、引号或空格异常:

print("训练集列名:", train_df.columns.tolist())
print("测试集列名:", test_df.columns.tolist())

2. 清理异常列名

如果输出的列名带有双引号(比如'"date"')或首尾空格,执行以下代码清理:

# 去除列名的双引号和首尾空格
train_df.columns = train_df.columns.str.strip('"').str.strip()
test_df.columns = test_df.columns.str.strip('"').str.strip()

这是因为你读文件时设置了quoting=2(即csv.QUOTE_ALL),会导致列名被双引号包裹,从而无法通过'date'正常访问。

3. 重新转换日期格式

列名清理完成后,再执行日期转换代码:

train_df['date'] = pd.to_datetime(train_df['date'], format='%b %d, %Y')
test_df['date'] = pd.to_datetime(test_df['date'], format='%b %d, %Y')

可选优化:调整读文件参数

后续读取文件时,可去掉quoting=2(或设置为quoting=0即csv.QUOTE_MINIMAL),避免列名被不必要的引号包裹:

train_df = pd.read_csv('/content/drugsComTrain_raw.tsv', sep='\t', quoting=0, dtype=dtypes)
test_df = pd.read_csv('/content/drugsComTest_raw.tsv', sep='\t', quoting=0, dtype=dtypes)

内容的提问来源于stack exchange,提问作者Karthik Bhandary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 10:02:12