You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

KeyError: 'date'错误求助——COVID-19数据季度可视化问题

COVID-19季度病例统计代码报错:KeyError: 'date'

问题背景

我正在用COVID-19数据集绘制疫情概览图表,尝试按季度分组计算每百万总病例数并用Seaborn可视化,但运行代码时触发KeyError: 'date'。奇怪的是date列并未缺失,且之前用该列实现的每日死亡数统计代码能正常运行,需要找出错误原因并解决。

出错代码

# Convert date column to datetime type
df['date'] = pd.to_datetime(df.date)

# Group data by quarter and calculate total cases per million
df_total_cases_quarterly = df.groupby(df['date'].dt.quarter).agg({'total_cases_per_million': 'sum'}).reset_index()

# Create bar plot for total cases per million on a quarterly basis
sns.set_style('whitegrid')
sns.barplot(x ='date', y = 'total_cases_per_million', data = df_total_cases_quarterly)
plt.title('Total Cases per Million (Worldwide) - Quarterly')
plt.xlabel('Quarter')
plt.ylabel('Total Cases per Million')
plt.xticks(ticks=[0, 1, 2, 3], labels=['Q1', 'Q2', 'Q3', 'Q4'])  # Update x-axis labels to show quarters
plt.show()

错误信息

---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
/opt/anaconda3/lib/python3.9/site-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
   3628             try:
-> 3629                 return self._engine.get_loc(casted_key)
   3630             except KeyError as err:

/opt/anaconda3/lib/python3.9/site-packages/pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()

/opt/anaconda3/lib/python3.9/site-packages/pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'date'

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
/var/folders/rf/8yc43r0d13l2gw8m9r17pc7h0000gn/T/ipykernel_942/1110610808.py in <module>
      3 
      4 # Group data by quarter and calculate total cases per million
----> 5 df_total_cases_quarterly = df.groupby(df['date'].dt.quarter).agg({'total_cases_per_million': 'sum'}).reset_index() # df['date'].dt.quarter
      6 
      7 # Create bar plot for total cases per million on a quarterly basis

/opt/anaconda3/lib/python3.9/site-packages/pandas/core/frame.py in __getitem__(self, key)
   3503             if self.columns.nlevels > 1:
   3504                 return self._getitem_multilevel(key)
-> 3505             indexer = self.columns.get_loc(key)
   3506             if is_integer(indexer):
   3507                 indexer = [indexer]

/opt/anaconda3/lib/python3.9/site-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
   3629                 return self._engine.get_loc(casted_key)
   3630             except KeyError as err:
-> 3631                 raise KeyError(key) from err
   3632             except TypeError:
   3633                 # If we have a listlike key, _check_indexing_error will raise

KeyError: 'date'

可正常运行的参考代码

# Filter data for daily new deaths per million
df_daily_new_deaths = df.groupby('date').agg({'new_deaths_per_million': 'sum'}).reset_index()

# Create line plot for daily new deaths per million
sns.set_style('whitegrid')
sns.lineplot(x = 'date', y = 'new_deaths_per_million', data = df_daily_new_deaths)
plt.title('Daily New Deaths per Million (Worldwide)')
plt.xlabel('Date')
plt.ylabel('Daily New Deaths per Million')
plt.xticks(rotation = 45)
plt.show()

原因分析与解决方案

错误原因

报错的核心是当前代码中的df变量已经丢失了date列。虽然你之前的代码能正常运行,但很可能在两段代码的执行过程中,df被意外修改了(比如被重新赋值为其他数据集、执行了删除列的操作),导致调用df['date']时找不到该列。

解决方案

  1. 检查当前df的结构
    先执行以下代码确认df的列情况:

    print(df.columns.tolist())
    

    如果输出里没有date,说明df已经被修改,需要重新加载原始数据集。

  2. 重新加载原始数据(关键修复)
    在运行季度统计代码前,确保使用的是完整的原始数据集,不要复用之前操作后被修改的df。比如重新执行数据加载代码:

    # 示例:重新加载数据,根据你的实际加载方式调整
    import pandas as pd
    df = pd.read_csv('your_covid_dataset.csv')
    df['date'] = pd.to_datetime(df['date'])
    
  3. 优化分组代码(避免后续问题)
    可以先给数据集添加明确的quarter列,再按该列分组,代码更清晰且不易出错:

    # 添加季度列
    df['quarter'] = df['date'].dt.quarter
    
    # 按季度分组计算
    df_total_cases_quarterly = df.groupby('quarter', as_index=False).agg(
        total_cases_per_million=('total_cases_per_million', 'sum')
    )
    
    # 绘图时使用quarter列作为x轴
    sns.set_style('whitegrid')
    sns.barplot(x='quarter', y='total_cases_per_million', data=df_total_cases_quarterly)
    plt.title('Total Cases per Million (Worldwide) - Quarterly')
    plt.xlabel('Quarter')
    plt.ylabel('Total Cases per Million')
    plt.xticks(ticks=[0, 1, 2, 3], labels=['Q1', 'Q2', 'Q3', 'Q4'])
    plt.show()
    

内容的提问来源于stack exchange,提问作者Taiwo Feyijimi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 09:25:06