如何在Python的statsmodels中忽略NaN?AnovaRM缺失值处理问题
Hey there, I’ve struggled with exactly this issue when working with repeated measures ANOVA in statsmodels, so I totally get your frustration! Let’s break down what’s going on and how to fix it:
Why missing='drop' doesn’t work for AnovaRM
Unlike OLS regression in statsmodels, AnovaRM does not support the missing='drop' parameter out of the box. This is because traditional repeated measures ANOVA relies on a balanced design—each subject (id in your code) needs complete data across all levels of the within-subjects factor (iv). When there are missing values, the function can’t compute the necessary variance components for the repeated measures, so it returns NaN for F-values and p-values instead of automatically dropping incomplete cases.
Practical Solutions
1. Switch to Mixed Linear Models (Recommended)
The most robust way to handle missing values in repeated measures designs is to use a mixed linear model instead of traditional ANOVA. Statsmodels' MixedLM supports unbalanced data and missing values by using likelihood-based estimation (ML/REML), which doesn’t require dropping entire subjects or observations.
Here’s how to adapt your code:
import statsmodels.api as sm from statsmodels.regression.mixed_linear_model import MixedLM from statsmodels.stats.anova import anova_lm # Fit the mixed model: RT predicted by iv, with random intercepts for each id model = MixedLM.from_formula( formula='RT ~ iv', groups='id', data=df3 ) result = model.fit() # Run ANOVA to get F-values and p-values for your within-subjects factor anova_results = anova_lm(result) print(anova_results)
This approach keeps all available data from each subject, even if they’re missing some conditions—no need to exclude entire samples.
2. Impute Missing Values First
If you specifically need to use AnovaRM, you can impute the missing values before running the analysis. This fills in gaps while preserving the balanced design required by traditional ANOVA. For more reliable results, use multiple imputation instead of simple mean imputation:
from statsmodels.imputation.mice import MICEData from statsmodels.stats.anova import AnovaRM # Create a MICE (Multiple Imputation by Chained Equations) object mice_data = MICEData(df3) # Generate an imputed dataset (you can also run multiple imputations and pool results) imputed_df = mice_data.data # Run AnovaRM on the imputed data aovrm = AnovaRM(imputed_df, 'RT', 'id', within=['iv']) anova_result = aovrm.fit() print(anova_result)
Note: Imputation introduces assumptions about your data—make sure to choose an imputation method that matches your variable type (e.g., mean imputation for continuous data like RT, or predictive mean matching).
3. Partial Data Exclusion (Last Resort)
If neither of the above works, you can exclude only the specific missing observations (not entire subjects) and use a univariate ANOVA for each condition, but this loses the repeated measures advantage (you can’t account for within-subject variance). This is only recommended if you have no other options.
内容的提问来源于stack exchange,提问作者knut_h

