R中lm模型报quantile.default缺失值错误,但数据集无缺失值求助
quantile.default(resid) Error in R's lm() for Linear Regression Hey there, let's dig into this error you're hitting with lm() in R. Even though you've confirmed no missing values in your dataset, the error about NaNs in residuals usually pops up because something goes wrong during model fitting—even if your raw data looks clean. Let's break down the most likely causes and fixes:
1. Check for Hidden Inf/-Inf Values in Derived Features
You mentioned creating log and squared features, which are common culprits here:
Logarithm features: If your original row/column counts include 0 or negative values (even if no missing values),
log(0)returns-Infandlog(negative)returnsNaN. These values get treated as missing bylm(), even if you didn't notice them at first.- Verify this with:
# Check for infinite values in all columns any(is.infinite(df)) # Check for NaNs specifically in derived log features any(is.na(df$log_rows) | is.infinite(df$log_rows)) - Fix: If you have 0s, use
log(x + 1)to avoid-Inf, or filter out rows where row/column counts are 0 if they don't make sense for your use case.
- Verify this with:
Squared features: While squaring rarely creates NaNs, extremely large values can overflow to
Infin R. Check for this withsummary(df)to see if any squared features have max values listed asInf.
2. Diagnose Multicollinearity Issues
If your derived features (like rows and rows^2, or log(rows) and rows) are highly correlated, this can cause the regression matrix to become singular. When this happens, lm() might fail to compute coefficients properly, leading to NaN residuals.
- Check correlation between predictors:
# Compute correlation matrix for all predictors (exclude 'time') cor_matrix <- cor(df[, !names(df) %in% "time"]) # Print pairs with correlation > 0.95 which(abs(cor_matrix) > 0.95 & cor_matrix != 1, arr.ind = TRUE) - Fix: Remove one of the highly correlated features, or use a regularized regression method like
glmnetinstead of plainlm()to handle multicollinearity.
3. Test a Simplified Model First
Your code snippet cuts off at time ~ ro...—make sure your formula is correctly written, and test with a minimal model to isolate the issue:
- Start with a basic model to confirm it works:
# Test simple model with only raw row/column counts simple_model <- lm(time ~ rows + columns, data = df) # Check if this runs without error summary(simple_model) - If this works, gradually add your derived features one by one (e.g., add
log_rows, thenrows_squared, etc.) until you hit the error. This will tell you exactly which feature is causing the problem.
4. Check for Extreme Outliers
Extreme values in your time variable or predictors can throw off the regression fit, leading to unexpected NaNs in residuals.
- Visualize outliers with boxplots:
boxplot(df$time, main = "Outliers in Training Time") boxplot(df$rows, main = "Outliers in Row Count") boxplot(df$columns, main = "Outliers in Column Count") - Fix: If outliers are due to data entry errors, remove them. If they're valid but extreme, consider transforming the
timevariable (e.g., log-transform) or using a robust regression method likerlm()from theMASSpackage.
内容的提问来源于stack exchange,提问作者Clock Slave

