You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中按条件筛选数据框:移除特定前缀且均值小于0.1的列

Solution to Remove Columns Starting with "x" with Mean < 0.1

First, let's recreate your data frame to work with:

# Create the original data frame
df <- data.frame(
  id = c(76978, 59911, 46537, 77345, 53180, 21063, 35456),
  name = c("phil", "jose", "matt", "benn", "crai", "lour", "moni"),
  class = c(2, 2, 3, 4, 2, 4, 4),
  x101 = c(0.407034783, 0.327173661, 0.590337464, 0.293847569, 0.844581456, 0.080756674, 0.445965164),
  x202 = c(0.001, 0.004, 0.005, 0.002, 0.003, 0.002, 0.004),
  x303 = c(0.192229687, 0.227843273, 0.057271545, 0.170405643, 0.253665748, 0.902143356, 0.531952568)
)

Base R Approach

Here's a straightforward way using base R functions:

  1. Flag columns starting with "x": Use grepl() to match column names that start with "x".
  2. Calculate column means: Get the average value for each of these x-columns.
  3. Identify columns to remove: Pick out the x-columns where the mean is less than 0.1.
  4. Clean the data frame: Remove those columns from the original data set.
# Step 1: Get logical vector of columns starting with "x"
x_cols <- grepl("^x", colnames(df))

# Step 2: Calculate means for x-columns
x_means <- colMeans(df[, x_cols])

# Step 3: Find columns to remove (mean < 0.1)
cols_to_remove <- names(x_means)[x_means < 0.1]

# Step 4: Remove the columns
df_cleaned_base <- df[, !colnames(df) %in% cols_to_remove]

Tidyverse (dplyr) Approach

If you prefer using the tidyverse ecosystem, here's a concise, readable method with dplyr:

library(dplyr)

# First, identify columns to remove
cols_to_remove <- df %>%
  select(starts_with("x")) %>%  # Keep only columns starting with "x"
  colMeans() %>%                # Calculate their average values
  keep(~ . < 0.1) %>%           # Filter those with mean < 0.1
  names()                       # Get the names of these columns

# Remove the unwanted columns from the original data frame
df_cleaned_tidy <- df %>%
  select(-all_of(cols_to_remove))

Result

Both methods will give you the same cleaned data frame. Only the x202 column is removed (its mean is 0.003, which is less than 0.1). Here's what the final output looks like:

print(df_cleaned_base)

Output:

id name class      x101      x303
1 76978 phil     2 0.4070348 0.1922297
2 59911 jose     2 0.3271737 0.2278433
3 46537 matt     3 0.5903375 0.0572715
4 77345 benn     4 0.2938476 0.1704056
5 53180 crai     2 0.8445815 0.2536657
6 21063 lour     4 0.0807567 0.9021434
7 35456 moni     4 0.4459652 0.5319526

内容的提问来源于stack exchange,提问作者GaRaGe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:21:50