You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中seasonal_decompose的resid丢失列名致Dickey-Fuller测试报错求助

解决Dickey-Fuller检验中KeyError: 'total_cases'的问题

嗨,我来帮你搞定这个报错!你遇到的KeyError本质是数据类型不匹配导致的:

  • 你一开始的indexedDataset_logScale是带列名total_cases的单列DataFrame
  • 但seasonal_decompose返回的resid是个pandas Series——它只有一个Name: resid的属性,没有列名索引,所以当你在test_stationarity函数里用timeseries['total_cases']去访问时,自然找不到这个键啦。

下面给你两种简单的修复方案,选哪个都能解决问题:


方案1:把residual Series转成带原列名的DataFrame

在处理完residual后,加一行代码把它转换成DataFrame并指定列名为total_cases,这样就能和原数据集的结构完全匹配:

from statsmodels.tsa.seasonal import seasonal_decompose
import numpy as np
import pandas as pd

indexedDataset_logScale.replace([np.inf, -np.inf], np.nan,inplace=True)
decomposition = seasonal_decompose(indexedDataset_logScale,freq=7)
residual = decomposition.resid
decomposedLogData = residual.dropna()

# 关键一步:将Series转为DataFrame,指定列名为total_cases
decomposedLogData = pd.DataFrame(decomposedLogData, columns=['total_cases'])

test_stationarity(decomposedLogData)

方案2:修改test_stationarity函数,兼容Series和DataFrame

调整函数逻辑,让它自动识别输入类型,不管传Series还是DataFrame都能正常跑:

from statsmodels.tsa.stattools import adfuller
import matplotlib.pyplot as plt

def test_stationarity(timeseries):
    # 处理兼容问题:如果是Series就转成DataFrame列,否则直接取原列
    if isinstance(timeseries, pd.Series):
        ts_data = timeseries.to_frame(name='total_cases')
    else:
        ts_data = timeseries['total_cases']
    
    movingAverage = ts_data.rolling(window=7).mean()
    movingStd = ts_data.rolling(window=7).std()
    
    orig = plt.plot(ts_data, color='blue', label='Origin')
    mean = plt.plot(movingAverage, color='red', label='Rolling Mean')
    std = plt.plot(movingStd, color='black', label='Rolling Std')
    plt.legend(loc='best')
    plt.title('Rolling Mean & Standard Deviation')
    plt.show(block = False)
    
    print('Result of Dickey-Fuller Test: ')
    # 直接用处理后的ts_data传入adfuller
    dftest = adfuller(ts_data, autolag='AIC')
    dfoutput = pd.Series(dftest[0:4], index=['Test Statistic', 'p-value', '#Lags Used', 'Number of Observations Used'])
    for key,value in dftest[4].items():
        dfoutput['Critical Value (%s)' %key] = value
    print(dfoutput)

改完之后,不管你传的是decomposedLogData(Series)还是原数据集(DataFrame),函数都能正常执行Dickey-Fuller检验。

额外小技巧

其实adfuller函数本身就支持直接传入Series或者一维数组,所以你也可以直接简化函数,去掉列名索引的部分,这样更省事:

# 简化版函数,直接兼容Series和DataFrame
def test_stationarity(timeseries):
    movingAverage = timeseries.rolling(window=7).mean()
    movingStd = timeseries.rolling(window=7).std()
    
    orig = plt.plot(timeseries, color='blue', label='Origin')
    mean = plt.plot(movingAverage, color='red', label='Rolling Mean')
    std = plt.plot(movingStd, color='black', label='Rolling Std')
    plt.legend(loc='best')
    plt.title('Rolling Mean & Standard Deviation')
    plt.show(block = False)
    
    print('Result of Dickey-Fuller Test: ')
    # 直接传入timeseries,adfuller会自动处理
    dftest = adfuller(timeseries, autolag='AIC')
    dfoutput = pd.Series(dftest[0:4], index=['Test Statistic', 'p-value', '#Lags Used', 'Number of Observations Used'])
    for key,value in dftest[4].items():
        dfoutput['Critical Value (%s)' %key] = value
    print(dfoutput)

这种情况下,你甚至不需要转换数据类型,直接把decomposedLogData(Series)传进去就能正常运行啦!


内容的提问来源于stack exchange,提问作者Hrithik Raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 07:07:39