You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中lm模型报quantile.default缺失值错误,但数据集无缺失值求助

Troubleshooting the quantile.default(resid) Error in R's lm() for Linear Regression

Hey there, let's dig into this error you're hitting with lm() in R. Even though you've confirmed no missing values in your dataset, the error about NaNs in residuals usually pops up because something goes wrong during model fitting—even if your raw data looks clean. Let's break down the most likely causes and fixes:

1. Check for Hidden Inf/-Inf Values in Derived Features

You mentioned creating log and squared features, which are common culprits here:

  • Logarithm features: If your original row/column counts include 0 or negative values (even if no missing values), log(0) returns -Inf and log(negative) returns NaN. These values get treated as missing by lm(), even if you didn't notice them at first.

    • Verify this with:
      # Check for infinite values in all columns
      any(is.infinite(df))
      # Check for NaNs specifically in derived log features
      any(is.na(df$log_rows) | is.infinite(df$log_rows))
      
    • Fix: If you have 0s, use log(x + 1) to avoid -Inf, or filter out rows where row/column counts are 0 if they don't make sense for your use case.
  • Squared features: While squaring rarely creates NaNs, extremely large values can overflow to Inf in R. Check for this with summary(df) to see if any squared features have max values listed as Inf.

2. Diagnose Multicollinearity Issues

If your derived features (like rows and rows^2, or log(rows) and rows) are highly correlated, this can cause the regression matrix to become singular. When this happens, lm() might fail to compute coefficients properly, leading to NaN residuals.

  • Check correlation between predictors:
    # Compute correlation matrix for all predictors (exclude 'time')
    cor_matrix <- cor(df[, !names(df) %in% "time"])
    # Print pairs with correlation > 0.95
    which(abs(cor_matrix) > 0.95 & cor_matrix != 1, arr.ind = TRUE)
    
  • Fix: Remove one of the highly correlated features, or use a regularized regression method like glmnet instead of plain lm() to handle multicollinearity.

3. Test a Simplified Model First

Your code snippet cuts off at time ~ ro...—make sure your formula is correctly written, and test with a minimal model to isolate the issue:

  • Start with a basic model to confirm it works:
    # Test simple model with only raw row/column counts
    simple_model <- lm(time ~ rows + columns, data = df)
    # Check if this runs without error
    summary(simple_model)
    
  • If this works, gradually add your derived features one by one (e.g., add log_rows, then rows_squared, etc.) until you hit the error. This will tell you exactly which feature is causing the problem.

4. Check for Extreme Outliers

Extreme values in your time variable or predictors can throw off the regression fit, leading to unexpected NaNs in residuals.

  • Visualize outliers with boxplots:
    boxplot(df$time, main = "Outliers in Training Time")
    boxplot(df$rows, main = "Outliers in Row Count")
    boxplot(df$columns, main = "Outliers in Column Count")
    
  • Fix: If outliers are due to data entry errors, remove them. If they're valid but extreme, consider transforming the time variable (e.g., log-transform) or using a robust regression method like rlm() from the MASS package.

内容的提问来源于stack exchange,提问作者Clock Slave

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:19:29