You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中Logistic回归内存不足:变量选择与数据集规模疑问

Hey there! Let's work through your problem together—first off, let's clear up a common misconception: 26 variables and 150k records shouldn't be causing a 130GB memory error on their own. That error is almost certainly not about the dataset size being "too big"—it's likely tied to how your variables are structured or the tools you're using. Let's break this down step by step:

First: Diagnose the Real Cause of the Memory Error

Before jumping into variable selection, let's fix the root memory issue:

  • Check for high-cardinality categorical variables: If you have a factor variable with thousands/millions of levels (like a unique household ID or full address that's been converted to a factor), R will create a massive dummy variable matrix that eats up memory. Run str(your_data) to inspect variable types—look for factors with an absurd number of levels. Fix this by either dropping the variable, using target/hash encoding for high-cardinality features, or keeping it as a character if it's not useful.
  • Clean up your R session: Run gc() to free up unused memory, and use ls() to check if there are other large objects (like raw datasets you've loaded but aren't using) cluttering up space.
  • Switch to a memory-efficient modeling function: The base glm() is solid, but packages like speedglm are optimized for large datasets and use far less memory. Give speedglm::speedglm(response ~ ., data = your_data, family = binomial()) a shot—it might resolve the memory issue without needing to drop variables.

Second: Variable Selection Without Running a Full Model

If you still want to narrow down variables (or need to for other reasons), you don't have to rely on step() which requires a full model fit. Here are better alternatives:

  • Regularized regression (Lasso): This is my top recommendation. The glmnet package does logistic regression with L1 regularization (Lasso), which automatically selects important variables while handling multicollinearity. It's designed for large datasets and doesn't require fitting a full model upfront. Example code:
    library(glmnet)
    # Create model matrix (drop intercept since glmnet adds it automatically)
    x <- model.matrix(response ~ ., data = your_data)[, -1]
    y <- your_data$response
    # Cross-validate to find the optimal regularization strength
    cv_fit <- cv.glmnet(x, y, family = "binomial")
    # View which variables were selected (non-zero coefficients)
    coef(cv_fit, s = "lambda.min")
    # Fit the final model with the best lambda
    final_model <- glmnet(x, y, family = "binomial", lambda = cv_fit$lambda.min)
    
  • Univariate screening (with caution): You can run single-variable logistic regressions for each predictor, then keep only those with statistically significant associations (adjust the p-value threshold to your needs). Just be aware this can miss variables that only matter in combination with others, and it doesn't account for multicollinearity. Example snippet:
    # Get p-values for each single-variable model
    var_pvals <- sapply(setdiff(names(your_data), "response"), function(var) {
      fit <- glm(response ~ ., data = your_data[, c("response", var)], family = "binomial")
      coef(summary(fit))[2, 4]
    })
    # Keep variables with p-value < 0.1 (adjust threshold as needed)
    selected_vars <- names(var_pvals)[var_pvals < 0.1]
    # Fit the reduced model
    reduced_fit <- glm(response ~ ., data = your_data[, c("response", selected_vars)], family = "binomial")
    
  • Subset selection tools: Packages like leaps use efficient algorithms to explore variable subsets without fitting every possible model, which is feasible with 26 variables (since 2^26 is a big number, but leaps prunes the search space). This is a middle ground between stepwise and regularization.

Final Note

Your 150k-record dataset is totally manageable in R—don't let the memory error make you think it's too large. The key is fixing variable structure issues and using tools optimized for large data instead of base glm + step().

内容的提问来源于stack exchange,提问作者statsnewbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:45:12