在R中按条件筛选数据框:移除特定前缀且均值小于0.1的列
Solution to Remove Columns Starting with "x" with Mean < 0.1
First, let's recreate your data frame to work with:
# Create the original data frame df <- data.frame( id = c(76978, 59911, 46537, 77345, 53180, 21063, 35456), name = c("phil", "jose", "matt", "benn", "crai", "lour", "moni"), class = c(2, 2, 3, 4, 2, 4, 4), x101 = c(0.407034783, 0.327173661, 0.590337464, 0.293847569, 0.844581456, 0.080756674, 0.445965164), x202 = c(0.001, 0.004, 0.005, 0.002, 0.003, 0.002, 0.004), x303 = c(0.192229687, 0.227843273, 0.057271545, 0.170405643, 0.253665748, 0.902143356, 0.531952568) )
Base R Approach
Here's a straightforward way using base R functions:
- Flag columns starting with "x": Use
grepl()to match column names that start with "x". - Calculate column means: Get the average value for each of these x-columns.
- Identify columns to remove: Pick out the x-columns where the mean is less than 0.1.
- Clean the data frame: Remove those columns from the original data set.
# Step 1: Get logical vector of columns starting with "x" x_cols <- grepl("^x", colnames(df)) # Step 2: Calculate means for x-columns x_means <- colMeans(df[, x_cols]) # Step 3: Find columns to remove (mean < 0.1) cols_to_remove <- names(x_means)[x_means < 0.1] # Step 4: Remove the columns df_cleaned_base <- df[, !colnames(df) %in% cols_to_remove]
Tidyverse (dplyr) Approach
If you prefer using the tidyverse ecosystem, here's a concise, readable method with dplyr:
library(dplyr) # First, identify columns to remove cols_to_remove <- df %>% select(starts_with("x")) %>% # Keep only columns starting with "x" colMeans() %>% # Calculate their average values keep(~ . < 0.1) %>% # Filter those with mean < 0.1 names() # Get the names of these columns # Remove the unwanted columns from the original data frame df_cleaned_tidy <- df %>% select(-all_of(cols_to_remove))
Result
Both methods will give you the same cleaned data frame. Only the x202 column is removed (its mean is 0.003, which is less than 0.1). Here's what the final output looks like:
print(df_cleaned_base)
Output:
id name class x101 x303 1 76978 phil 2 0.4070348 0.1922297 2 59911 jose 2 0.3271737 0.2278433 3 46537 matt 3 0.5903375 0.0572715 4 77345 benn 4 0.2938476 0.1704056 5 53180 crai 2 0.8445815 0.2536657 6 21063 lour 4 0.0807567 0.9021434 7 35456 moni 4 0.4459652 0.5319526
内容的提问来源于stack exchange,提问作者GaRaGe
相关产品推荐
相关产品推荐

