You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

时间序列NaN填充规则及statsmodels ADF检验缺失值处理问询

Got it, let's tackle your two time series data handling needs with practical, actionable code and explanations using Python's pandas and statsmodels libraries.

1. NaN Handling for Time Series DataFrames

Your rule is clear: if the first observation is NaN, fill all leading NaNs with the first valid value; otherwise, fill subsequent NaNs with the previous observation. Here's how to implement this cleanly:

import pandas as pd

def custom_fill_nan(col):
    # Get the index of the first valid value in the column
    first_valid_idx = col.first_valid_index()
    
    # If the first row is NaN (first_valid_idx isn't the first position)
    if first_valid_idx != 0:
        # Fill all leading NaNs with the first valid value
        col.loc[:first_valid_idx] = col[first_valid_idx]
    
    # Fill remaining NaNs with the previous valid observation (forward fill)
    return col.fillna(method='ffill')

# Apply the function to every column in your DataFrame
processed_df = df.apply(custom_fill_nan, axis=0)

How this works:

  • We first target leading NaNs specifically: if the first row is missing, we grab the first real value and backfill all empty spots before it.
  • For any remaining NaNs scattered later in the sequence, we use pandas' ffill (forward fill) to carry over the last valid observation, which aligns with your "沿用前一观测值" rule.
2. ADF Test Workflow with NaN Filtering & Handling

For your statsmodels ADF test scenario, we need to first filter out sequences with too many NaNs, then clean the remaining ones to work with adfuller() (which can't handle missing values). Here's a step-by-step implementation:

from statsmodels.tsa.stattools import adfuller

# Define your NaN threshold (e.g., 10% missing values allowed)
nan_threshold = 0.1

# Store ADF results for valid sequences
adf_results = {}

for col_name, col_data in df.items():
    # Calculate NaN ratio for the column
    nan_ratio = col_data.isna().mean()
    
    # Skip columns with too many NaNs
    if nan_ratio > nan_threshold:
        print(f"Skipping {col_name}: NaN ratio ({nan_ratio:.2f}) exceeds threshold {nan_threshold}")
        continue
    
    # Step 1: Handle first-row NaN
    first_valid_idx = col_data.first_valid_index()
    cleaned_col = col_data.copy()
    if pd.isna(cleaned_col.iloc[0]):
        # Replace first NaN with the first valid value in the sequence
        cleaned_col.iloc[0] = cleaned_col[first_valid_idx]
    
    # Step 2: Fill remaining NaNs with previous observation
    cleaned_col = cleaned_col.fillna(method='ffill')
    
    # Run ADF test
    adf_result = adfuller(cleaned_col)
    adf_results[col_name] = {
        'ADF Statistic': adf_result[0],
        'p-value': adf_result[1],
        'Critical Values': adf_result[4]
    }

# Print results for all valid sequences
for col, res in adf_results.items():
    print(f"\nADF Test Results for {col}:")
    print(f"ADF Statistic: {res['ADF Statistic']:.4f}")
    print(f"p-value: {res['p-value']:.4f}")
    print("Critical Values:")
    for key, val in res['Critical Values'].items():
        print(f"  {key}: {val:.4f}")

Key Notes:

  • NaN Filtering: We calculate the proportion of missing values per column; if it's over your threshold c, we skip the column entirely to avoid biased results.
  • Missing Value Cleaning: We strictly follow your rules: first fix the initial NaN with the first valid value, then forward-fill any remaining gaps. This ensures the sequence is continuous, which is required for the ADF test.
  • ADF Results: We store the key outputs (statistic, p-value, critical values) in a dictionary for easy analysis—you can adjust this to save to a DataFrame or file if needed.

内容的提问来源于stack exchange,提问作者eternity1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:57:00