时间序列Dickey-Fuller测试报错:ValueError: too many values to unpack (expected 2)
ValueError: too many values to unpack in Dickey-Fuller Test Hey there! Let's break down exactly what's going on with your Dickey-Fuller test error, and clear up those confusing dftest index values once and for all.
First: What does adfuller() actually return?
I’m assuming you’re using the adfuller() function from statsmodels.tsa.stattools (the standard tool for this test). This function returns a 6-element tuple with specific values in order:
dftest[0]: The test statistic (the core number from the DF test)dftest[1]: The p-value (the key metric to judge stationarity)dftest[2]: The number of lags used in the test (to account for autocorrelation)dftest[3]: The number of valid observations included in the testdftest[4]: A dictionary of critical values (for 1%, 5%, and 10% significance levels)dftest[5]: The maximum information criterion (like AIC, used to pick optimal lags)
Why you’re getting the too many values to unpack error
Your error almost certainly comes from trying to unpack a subset of this tuple into fewer variables than elements exist. For example:
If your code looks like this:
stat, p_val = dftest[0:4] # dftest[0:4] gives 4 values, but you're trying to put them into 2 variables
That’s a problem—4 values can’t fit into 2 variables, hence the ValueError.
How to fix it
You have two clean ways to handle this:
Option 1: Unpack all values at once (most readable)
from statsmodels.tsa.stattools import adfuller # Run the test on your cleaned, NaN-free data dftest = adfuller(df['your_time_series_column']) # Unpack all 6 values explicitly test_stat, p_val, used_lags, n_obs, crit_vals, max_ic = dftest
Option 2: Index directly to get specific values
If you only need certain metrics, call them by their index without unpacking:
# Get just the test statistic and p-value test_stat = dftest[0] p_val = dftest[1] # Get the critical values (e.g., 5% level) critical_value_5pct = dftest[4]['5%']
Quick guide to interpreting each value
dftest[0](test statistic): The larger the absolute value, the stronger the evidence against the "non-stationary" null hypothesis.dftest[1](p-value): If this is less than 0.05 (your typical significance level), you can reject the null hypothesis and conclude your data is stationary.dftest[4](critical values): If your test statistic is smaller than the critical value for your chosen significance level, you also reject the null hypothesis.
Full working example
Here’s a complete snippet to adapt to your data:
import pandas as pd from statsmodels.tsa.stattools import adfuller # Load and clean your data (you already did the NaN removal step) df = pd.read_csv('your_data.csv') clean_df = df.dropna(subset=['your_time_series_column']) # Run ADF test dftest = adfuller(clean_df['your_time_series_column']) # Print results clearly print(f"ADF Test Statistic: {dftest[0]:.4f}") print(f"P-Value: {dftest[1]:.4f}") print(f"Lags Used: {dftest[2]}") print(f"Critical Values: {dftest[4]}") # Make a stationarity judgment if dftest[1] < 0.05: print("\n✅ Data is stationary (reject null hypothesis)") else: print("\n❌ Data is non-stationary (fail to reject null hypothesis)")
内容的提问来源于stack exchange,提问作者Josh

