You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

不使用caret包如何移除近零方差?附新手待修改代码

Hey there! Let's tackle this together since you're new to programming and want to avoid using caret::nearZeroVar(). First, let's break down what that function does: it identifies features (columns in your matrix X) with near-zero variance—meaning either almost all values are identical, or the variance is so low that the feature doesn't add any meaningful signal to your model. Here's how you can replicate this logic with base R, no extra packages needed:

1. Core Logic Behind nearZeroVar

The function typically uses two key checks to flag columns:

  • The percentage of unique values relative to total samples (e.g., if less than 10% of values are unique)
  • The ratio of the most frequent value to the second most frequent (e.g., if the top value is 20x more common than the runner-up)

2. Manual Implementation for Your Dataset

Since your X is a matrix, we can use vectorized operations to scan each column and apply these checks:

# Define adjustable thresholds (tweak these based on your needs)
unique_pct_threshold <- 0.1  # Require at least 10% unique values
freq_ratio_threshold <- 20    # Max allowed ratio of top two value frequencies

# Helper function to check if a single column has near-zero variance
check_near_zero <- function(col) {
  value_counts <- table(col)
  # Case 1: All values are identical
  if (length(value_counts) == 1) {
    return(TRUE)
  }
  # Case 2: Too few unique values
  unique_pct <- length(value_counts) / length(col)
  # Case 3: Top value is way more frequent than second
  sorted_counts <- sort(value_counts, decreasing = TRUE)
  freq_ratio <- sorted_counts[1] / sorted_counts[2]
  
  # Return TRUE if any condition is met
  return(unique_pct < unique_pct_threshold || freq_ratio > freq_ratio_threshold)
}

# Apply check to all columns in X
nzv_columns <- apply(X, 2, check_near_zero)

# Remove the near-zero variance columns from your matrix
X_clean <- X[, !nzv_columns]

3. Simplified Option: Remove Exactly Zero Variance Columns

If you only need to get rid of columns with zero variance (all values the same), you can use a much shorter approach:

# Calculate variance for each column
column_variances <- apply(X, 2, var)

# Keep only columns with variance > 0
X_clean <- X[, column_variances > 0]

Quick Note on Your Code

Looking at your snippet, you're building a 4-class dataset with one-hot encoded labels in Y. After cleaning X, just double-check that the number of rows in X_clean still matches Y (it will, since we're only removing columns, not rows) and you're good to proceed with your modeling.

内容的提问来源于stack exchange,提问作者Bad Attitude

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:42:23