You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于波士顿数据集的Logistic Regression输出异常问题咨询

Hey there! Let's dig into why your logistic regression output might be off with the Boston dataset. First, let's break down potential issues in your current code and actionable steps to fix them:

Key Issues to Check & Fix

1. The attach() call is likely causing confusion

When you run attach(Boston), you're linking the original Boston dataset to your R session's search path—but you already created ourBoston which includes your new high.medv factor. This can lead to naming conflicts or accidentally pulling data from the wrong dataset (e.g., using the original medv instead of the columns in ourBoston).

Fix: Ditch attach() entirely and use explicit dataframe referencing (like ourBoston$column_name) for all operations. It's cleaner and avoids weird bugs.

2. Missing (or incorrect) train/test split code

You mentioned a 70% split but didn't include the code for it. If you're training on the full dataset without splitting, or if you didn't stratify the split to preserve the balance of high.medv classes, you could get misleading results (especially if one class is way more common than the other).

Example stratified split code (to keep class balance):

set.seed(123) # For reproducibility
library(caret)
train_idx <- createDataPartition(ourBoston$high.medv, p = 0.7, list = FALSE)
train_data <- ourBoston[train_idx, ]
test_data <- ourBoston[-train_idx, ]

3. Incomplete logistic regression model code

You haven't shown the actual glm() call that trains your model—this is where most mistakes happen! Common missteps here include:

  • Forgetting to specify family = binomial (required for logistic regression; without this, you're running linear regression instead)
  • Using an incorrect formula (e.g., missing predictors, misreferencing high.medv)

Example correct model call:

logit_model <- glm(high.medv ~ ., data = train_data, family = binomial)
summary(logit_model)

4. Class imbalance in high.medv

If most of your observations are "N" (or "Y"), the model might just predict the majority class to get high accuracy, but fail at actually learning the relationship. Let's check the balance first:

# Check class counts and proportions
table(ourBoston$high.medv)
prop.table(table(ourBoston$high.medv))

If imbalance is severe, try:

  • Oversampling the minority class
  • Undersampling the majority class
  • Adding class weights to your model:
    # Adjust the weight value based on your imbalance ratio
    logit_weighted <- glm(high.medv ~ ., data = train_data, family = binomial,
                          weights = ifelse(high.medv == "Y", 4, 1))
    

5. Multicollinearity among predictors

The Boston dataset has several highly correlated predictors (e.g., rm and medv, tax and rad). Multicollinearity inflates standard errors and makes coefficient interpretations unreliable. Check for this using Variance Inflation Factor (VIF):

library(car)
vif(logit_model)

If any VIF value is above 5-10, consider removing one of the correlated predictors or using regularization (like LASSO logistic regression with the glmnet package).

If you try these steps and still get wonky output, share the full model summary and any error messages you're seeing—we can dive deeper!

内容的提问来源于stack exchange,提问作者elle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:40:53