时间序列NaN填充规则及statsmodels ADF检验缺失值处理问询
Got it, let's tackle your two time series data handling needs with practical, actionable code and explanations using Python's pandas and statsmodels libraries.
1. NaN Handling for Time Series DataFrames
Your rule is clear: if the first observation is NaN, fill all leading NaNs with the first valid value; otherwise, fill subsequent NaNs with the previous observation. Here's how to implement this cleanly:
import pandas as pd def custom_fill_nan(col): # Get the index of the first valid value in the column first_valid_idx = col.first_valid_index() # If the first row is NaN (first_valid_idx isn't the first position) if first_valid_idx != 0: # Fill all leading NaNs with the first valid value col.loc[:first_valid_idx] = col[first_valid_idx] # Fill remaining NaNs with the previous valid observation (forward fill) return col.fillna(method='ffill') # Apply the function to every column in your DataFrame processed_df = df.apply(custom_fill_nan, axis=0)
How this works:
- We first target leading NaNs specifically: if the first row is missing, we grab the first real value and backfill all empty spots before it.
- For any remaining NaNs scattered later in the sequence, we use pandas'
ffill(forward fill) to carry over the last valid observation, which aligns with your "沿用前一观测值" rule.
2. ADF Test Workflow with NaN Filtering & Handling
For your statsmodels ADF test scenario, we need to first filter out sequences with too many NaNs, then clean the remaining ones to work with adfuller() (which can't handle missing values). Here's a step-by-step implementation:
from statsmodels.tsa.stattools import adfuller # Define your NaN threshold (e.g., 10% missing values allowed) nan_threshold = 0.1 # Store ADF results for valid sequences adf_results = {} for col_name, col_data in df.items(): # Calculate NaN ratio for the column nan_ratio = col_data.isna().mean() # Skip columns with too many NaNs if nan_ratio > nan_threshold: print(f"Skipping {col_name}: NaN ratio ({nan_ratio:.2f}) exceeds threshold {nan_threshold}") continue # Step 1: Handle first-row NaN first_valid_idx = col_data.first_valid_index() cleaned_col = col_data.copy() if pd.isna(cleaned_col.iloc[0]): # Replace first NaN with the first valid value in the sequence cleaned_col.iloc[0] = cleaned_col[first_valid_idx] # Step 2: Fill remaining NaNs with previous observation cleaned_col = cleaned_col.fillna(method='ffill') # Run ADF test adf_result = adfuller(cleaned_col) adf_results[col_name] = { 'ADF Statistic': adf_result[0], 'p-value': adf_result[1], 'Critical Values': adf_result[4] } # Print results for all valid sequences for col, res in adf_results.items(): print(f"\nADF Test Results for {col}:") print(f"ADF Statistic: {res['ADF Statistic']:.4f}") print(f"p-value: {res['p-value']:.4f}") print("Critical Values:") for key, val in res['Critical Values'].items(): print(f" {key}: {val:.4f}")
Key Notes:
- NaN Filtering: We calculate the proportion of missing values per column; if it's over your threshold
c, we skip the column entirely to avoid biased results. - Missing Value Cleaning: We strictly follow your rules: first fix the initial NaN with the first valid value, then forward-fill any remaining gaps. This ensures the sequence is continuous, which is required for the ADF test.
- ADF Results: We store the key outputs (statistic, p-value, critical values) in a dictionary for easy analysis—you can adjust this to save to a DataFrame or file if needed.
内容的提问来源于stack exchange,提问作者eternity1
相关产品推荐
相关产品推荐

