Python中seasonal_decompose的resid丢失列名致Dickey-Fuller测试报错求助
解决Dickey-Fuller检验中KeyError: 'total_cases'的问题
嗨,我来帮你搞定这个报错!你遇到的KeyError本质是数据类型不匹配导致的:
- 你一开始的
indexedDataset_logScale是带列名total_cases的单列DataFrame - 但
seasonal_decompose返回的resid是个pandas Series——它只有一个Name: resid的属性,没有列名索引,所以当你在test_stationarity函数里用timeseries['total_cases']去访问时,自然找不到这个键啦。
下面给你两种简单的修复方案,选哪个都能解决问题:
方案1:把residual Series转成带原列名的DataFrame
在处理完residual后,加一行代码把它转换成DataFrame并指定列名为total_cases,这样就能和原数据集的结构完全匹配:
from statsmodels.tsa.seasonal import seasonal_decompose import numpy as np import pandas as pd indexedDataset_logScale.replace([np.inf, -np.inf], np.nan,inplace=True) decomposition = seasonal_decompose(indexedDataset_logScale,freq=7) residual = decomposition.resid decomposedLogData = residual.dropna() # 关键一步:将Series转为DataFrame,指定列名为total_cases decomposedLogData = pd.DataFrame(decomposedLogData, columns=['total_cases']) test_stationarity(decomposedLogData)
方案2:修改test_stationarity函数,兼容Series和DataFrame
调整函数逻辑,让它自动识别输入类型,不管传Series还是DataFrame都能正常跑:
from statsmodels.tsa.stattools import adfuller import matplotlib.pyplot as plt def test_stationarity(timeseries): # 处理兼容问题:如果是Series就转成DataFrame列,否则直接取原列 if isinstance(timeseries, pd.Series): ts_data = timeseries.to_frame(name='total_cases') else: ts_data = timeseries['total_cases'] movingAverage = ts_data.rolling(window=7).mean() movingStd = ts_data.rolling(window=7).std() orig = plt.plot(ts_data, color='blue', label='Origin') mean = plt.plot(movingAverage, color='red', label='Rolling Mean') std = plt.plot(movingStd, color='black', label='Rolling Std') plt.legend(loc='best') plt.title('Rolling Mean & Standard Deviation') plt.show(block = False) print('Result of Dickey-Fuller Test: ') # 直接用处理后的ts_data传入adfuller dftest = adfuller(ts_data, autolag='AIC') dfoutput = pd.Series(dftest[0:4], index=['Test Statistic', 'p-value', '#Lags Used', 'Number of Observations Used']) for key,value in dftest[4].items(): dfoutput['Critical Value (%s)' %key] = value print(dfoutput)
改完之后,不管你传的是decomposedLogData(Series)还是原数据集(DataFrame),函数都能正常执行Dickey-Fuller检验。
额外小技巧
其实adfuller函数本身就支持直接传入Series或者一维数组,所以你也可以直接简化函数,去掉列名索引的部分,这样更省事:
# 简化版函数,直接兼容Series和DataFrame def test_stationarity(timeseries): movingAverage = timeseries.rolling(window=7).mean() movingStd = timeseries.rolling(window=7).std() orig = plt.plot(timeseries, color='blue', label='Origin') mean = plt.plot(movingAverage, color='red', label='Rolling Mean') std = plt.plot(movingStd, color='black', label='Rolling Std') plt.legend(loc='best') plt.title('Rolling Mean & Standard Deviation') plt.show(block = False) print('Result of Dickey-Fuller Test: ') # 直接传入timeseries,adfuller会自动处理 dftest = adfuller(timeseries, autolag='AIC') dfoutput = pd.Series(dftest[0:4], index=['Test Statistic', 'p-value', '#Lags Used', 'Number of Observations Used']) for key,value in dftest[4].items(): dfoutput['Critical Value (%s)' %key] = value print(dfoutput)
这种情况下,你甚至不需要转换数据类型,直接把decomposedLogData(Series)传进去就能正常运行啦!
内容的提问来源于stack exchange,提问作者Hrithik Raj
相关产品推荐
相关产品推荐

