You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中实现70:30数据拆分并仅在训练样本中引入异常值与多重共线性?

Solution: Contaminate Only Training Data (Keep Test Set Clean)

Got it, let's fix this step by step. The core problem with your current code is that you're adding multicollinearity and outliers to the entire dataset before splitting into train/test, which means your test set isn't "pure" as required. We need to reverse this workflow:

  1. Generate a clean, baseline dataset with no multicollinearity or outliers
  2. Split the clean data into training (70%) and test (30%) subsets first
  3. Introduce multicollinearity only to the training set's features
  4. Add outliers only to the training set's features and target
  5. Train models on the contaminated training set, then evaluate on the pure test set

Here's the revised code with these changes, plus explanations of key tweaks:

# Revised function: only contaminates training set, keeps test set clean
um <- function(R, n, sig, p, po, py, fx, fy){ 
  #' R: multicollinearity level (0-1)
  #' n: total sample size
  #' sig: error variance
  #' p: number of predictors
  #' po: proportion of x-outliers in training set
  #' py: proportion of y-outliers in training set
  #' fx: magnitude of x-outliers
  #' fy: magnitude of y-outliers
  
  RR <- 1000 
  set.seed(123) 
  OP1 <- NULL 
  
  for (i in 1:RR){ 
    # --------------------------
    # Step 1: Generate CLEAN base data (no multicollinearity, no outliers)
    # --------------------------
    # Clean predictors (no multicollinearity)
    x_raw <- matrix(rnorm(n*p, mean=0, sd=1), nrow=n, ncol=p) 
    # Clean target (no outliers)
    u <- rnorm(n, 0, sig) 
    # True coefficients (use first eigenvector of clean x, same as original logic)
    b <- eigen(t(x_raw)%*%x_raw)$vec[,1] 
    y_raw <- x_raw %*% b + u 
    
    # --------------------------
    # Step 2: Split clean data into train/test FIRST
    # --------------------------
    training_idx <- sample(1:n, size=round(n*0.7), replace=FALSE) 
    test_idx <- setdiff(1:n, training_idx) 
    
    # Extract clean test set (STAYS PURE!)
    x_test_clean <- x_raw[test_idx, ]
    y_test_clean <- y_raw[test_idx]
    
    # Extract clean training set (we'll contaminate this)
    x_train_clean <- x_raw[training_idx, ]
    y_train_clean <- y_raw[training_idx]
    n_train <- length(training_idx)
    
    # --------------------------
    # Step 3: Add multicollinearity ONLY to training predictors
    # --------------------------
    # Use the same logic as original, but only for training data
    W_train <- matrix(rnorm(n_train*(p+1), mean=0, sd=1), n_train, p+1) 
    x_train_contaminated <- matrix(0, nrow=n_train, ncol=p)
    for (j in 1:p){ 
      x_train_contaminated[,j] <- sqrt(1-R^2)*W_train[,j] + R*W_train[,p+1] 
    }
    
    # --------------------------
    # Step 4: Add outliers ONLY to training set
    # --------------------------
    # X-outliers (only in training set)
    rep1 <- sample(1:n_train, size=round(po*n_train), replace=FALSE) 
    x_train_contaminated[rep1, 2] <- fx*max(x_train_contaminated[,2]) + x_train_contaminated[rep1, 2] 
    
    # Y-outliers (only in training set)
    rep2 <- sample(1:n_train, size=round(py*n_train), replace=FALSE) 
    y_train_contaminated <- y_train_clean
    y_train_contaminated[rep2] <- fy*max(y_train_contaminated) + y_train_contaminated[rep2]
    
    # --------------------------
    # Step 5: Standardize correctly (avoid data leakage!)
    # --------------------------
    # Scale training data, then apply same scaling to test data
    # Scale predictors
    x_train_scaled <- scale(x_train_contaminated)
    x_test_scaled <- scale(x_test_clean, center=attr(x_train_scaled, "scaled:center"), scale=attr(x_train_scaled, "scaled:scale"))
    # Scale target
    y_train_scaled <- scale(y_train_contaminated)
    y_test_scaled <- scale(y_test_clean, center=attr(y_train_scaled, "scaled:center"), scale=attr(y_train_scaled, "scaled:scale"))
    
    # --------------------------
    # Step 6: Train models and evaluate on pure test set
    # --------------------------
    # Convert to matrices for modeling
    xtr <- cbind(1, x_train_scaled) # add intercept
    ytr <- y_train_scaled
    xte <- cbind(1, x_test_scaled)
    yte <- y_test_scaled
    
    # Train models
    mest <- rlm(ytr ~ xtr - 1, psi=psi.huber, k2=1.345, maxit=1000)$coefficients # remove intercept since we added it manually
    ols <- lm(ytr ~ xtr - 1)$coefficients
    
    # Calculate MdAE on pure test set
    OLS <- median(abs(yte - xte %*% ols))
    M <- median(abs(yte - xte %*% mest))
    res2 <- cbind(OLS, M) 
    OP1 <- rbind(OP1, res2)
  } 
  
  # Aggregate results
  MAE <- colMeans(OP1) # average over RR iterations
  data.frame(R, n, sig, p, po, py, fx, fy, OLS_MdAE=MAE[1], Robust_MdAE=MAE[2]) 
} 

# Run simulations
results <- NULL 
R <- c(0.99) 
n <- c(100) 
sig <- c(5) 
p <- c(5) 
po <- c(0.2) 
py <- c(0.2) 
fx <- c(5) 
fy <- c(5) 

for(i in 1:length(R)){ 
  for(j in 1:length(n)){ 
    for(k in 1:length(sig)){ 
      for(l in 1:length(p)){ 
        for(m in 1:length(po)){ 
          for(nn in 1:length(py)){ 
            for(o in 1:length(fx)){ 
              for(pp in 1:length(fy)){ 
                results <- rbind(results, um(R=R[i], n=n[j], sig=sig[k], p=p[l], po=po[m], py=py[nn], fx=fx[o], fy=fy[pp])) 
              } 
            } 
          } 
        } 
      } 
    } 
  } 
} 

View(results)

Key Changes Explained

  • Early Train/Test Split: We split the clean data before any contamination, so the test set remains completely untouched (no multicollinearity, no outliers).
  • Targeted Contamination: Multicollinearity and outliers are only applied to the training subset, not the entire dataset.
  • Correct Standardization: We scale the training data first, then apply the same center/scale values to the test set. This avoids data leakage (using test set statistics to scale training data).
  • Fixed RLM/LM Syntax: Removed the redundant intercept term since we manually added a column of 1s to the predictor matrices (avoids model fitting errors).
  • Aggregated Results: Averaged MdAE over all 1000 iterations for more stable estimates.

内容的提问来源于stack exchange,提问作者jeza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 20:57:52