KeyError: 'date'错误求助——COVID-19数据季度可视化问题
COVID-19季度病例统计代码报错:KeyError: 'date'
问题背景
我正在用COVID-19数据集绘制疫情概览图表,尝试按季度分组计算每百万总病例数并用Seaborn可视化,但运行代码时触发KeyError: 'date'。奇怪的是date列并未缺失,且之前用该列实现的每日死亡数统计代码能正常运行,需要找出错误原因并解决。
出错代码
# Convert date column to datetime type df['date'] = pd.to_datetime(df.date) # Group data by quarter and calculate total cases per million df_total_cases_quarterly = df.groupby(df['date'].dt.quarter).agg({'total_cases_per_million': 'sum'}).reset_index() # Create bar plot for total cases per million on a quarterly basis sns.set_style('whitegrid') sns.barplot(x ='date', y = 'total_cases_per_million', data = df_total_cases_quarterly) plt.title('Total Cases per Million (Worldwide) - Quarterly') plt.xlabel('Quarter') plt.ylabel('Total Cases per Million') plt.xticks(ticks=[0, 1, 2, 3], labels=['Q1', 'Q2', 'Q3', 'Q4']) # Update x-axis labels to show quarters plt.show()
错误信息
--------------------------------------------------------------------------- KeyError Traceback (most recent call last) /opt/anaconda3/lib/python3.9/site-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3628 try: -> 3629 return self._engine.get_loc(casted_key) 3630 except KeyError as err: /opt/anaconda3/lib/python3.9/site-packages/pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc() /opt/anaconda3/lib/python3.9/site-packages/pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc() pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item() pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item() KeyError: 'date' The above exception was the direct cause of the following exception: KeyError Traceback (most recent call last) /var/folders/rf/8yc43r0d13l2gw8m9r17pc7h0000gn/T/ipykernel_942/1110610808.py in <module> 3 4 # Group data by quarter and calculate total cases per million ----> 5 df_total_cases_quarterly = df.groupby(df['date'].dt.quarter).agg({'total_cases_per_million': 'sum'}).reset_index() # df['date'].dt.quarter 6 7 # Create bar plot for total cases per million on a quarterly basis /opt/anaconda3/lib/python3.9/site-packages/pandas/core/frame.py in __getitem__(self, key) 3503 if self.columns.nlevels > 1: 3504 return self._getitem_multilevel(key) -> 3505 indexer = self.columns.get_loc(key) 3506 if is_integer(indexer): 3507 indexer = [indexer] /opt/anaconda3/lib/python3.9/site-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3629 return self._engine.get_loc(casted_key) 3630 except KeyError as err: -> 3631 raise KeyError(key) from err 3632 except TypeError: 3633 # If we have a listlike key, _check_indexing_error will raise KeyError: 'date'
可正常运行的参考代码
# Filter data for daily new deaths per million df_daily_new_deaths = df.groupby('date').agg({'new_deaths_per_million': 'sum'}).reset_index() # Create line plot for daily new deaths per million sns.set_style('whitegrid') sns.lineplot(x = 'date', y = 'new_deaths_per_million', data = df_daily_new_deaths) plt.title('Daily New Deaths per Million (Worldwide)') plt.xlabel('Date') plt.ylabel('Daily New Deaths per Million') plt.xticks(rotation = 45) plt.show()
原因分析与解决方案
错误原因
报错的核心是当前代码中的df变量已经丢失了date列。虽然你之前的代码能正常运行,但很可能在两段代码的执行过程中,df被意外修改了(比如被重新赋值为其他数据集、执行了删除列的操作),导致调用df['date']时找不到该列。
解决方案
检查当前
df的结构
先执行以下代码确认df的列情况:print(df.columns.tolist())如果输出里没有
date,说明df已经被修改,需要重新加载原始数据集。重新加载原始数据(关键修复)
在运行季度统计代码前,确保使用的是完整的原始数据集,不要复用之前操作后被修改的df。比如重新执行数据加载代码:# 示例:重新加载数据,根据你的实际加载方式调整 import pandas as pd df = pd.read_csv('your_covid_dataset.csv') df['date'] = pd.to_datetime(df['date'])优化分组代码(避免后续问题)
可以先给数据集添加明确的quarter列,再按该列分组,代码更清晰且不易出错:# 添加季度列 df['quarter'] = df['date'].dt.quarter # 按季度分组计算 df_total_cases_quarterly = df.groupby('quarter', as_index=False).agg( total_cases_per_million=('total_cases_per_million', 'sum') ) # 绘图时使用quarter列作为x轴 sns.set_style('whitegrid') sns.barplot(x='quarter', y='total_cases_per_million', data=df_total_cases_quarterly) plt.title('Total Cases per Million (Worldwide) - Quarterly') plt.xlabel('Quarter') plt.ylabel('Total Cases per Million') plt.xticks(ticks=[0, 1, 2, 3], labels=['Q1', 'Q2', 'Q3', 'Q4']) plt.show()
内容的提问来源于stack exchange,提问作者Taiwo Feyijimi
相关产品推荐
相关产品推荐

