如何批量处理数据框,提取所有特征列的mean并移除括号内数据?
Got it, let's solve this efficiently—no need to repeat code for 109 columns! Your goal is to pull out the mean values from every feature column (each containing both mean and sd, formatted like mean_value (sd_value)), and we can do this in bulk using tidyverse tools in R.
Step 1: Load Required Packages
First, make sure you have the tidyverse installed and loaded—it combines dplyr for data manipulation and tidyr for reshaping data, which is perfect for this task:
library(tidyverse)
Step 2: Batch Extract Mean Values (Two Simple Methods)
Method 1: Using across() + str_extract()
This method uses regex to directly pull the mean value (everything before the first left parenthesis) from every target column, then converts it to a numeric value.
Suppose your feature columns have a consistent prefix like feature_ (adjust the selector if your columns are named differently—use everything() if all non-ID columns are features):
# Extract mean values and create new columns with "_mean" suffix df_mean_only <- df %>% mutate( across( .cols = starts_with("feature"), # Select all feature columns (adjust this to match your column names) .fns = ~ str_extract(., "^[^(]+") %>% trimws() %>% as.numeric(), .names = "{col}_mean" ) ) %>% # Optional: Drop the original feature columns if you don't need them select(-starts_with("feature"))
^[^(]+is the regex that grabs all characters from the start of the string up to (but not including) the first(.trimws()removes any extra spaces around the mean value (in case your formatting has gaps)..names = "{col}_mean"ensures each new column keeps the original feature name plus a_meansuffix.
Method 2: Using separate_wider_delim() (More Intuitive for Splitting)
If you want to temporarily split columns into both mean and sd (then just keep the mean), separate_wider_delim is a cleaner option for bulk splitting:
df_mean_only <- df %>% separate_wider_delim( cols = starts_with("feature"), # Target feature columns delim = " (", # The separator between mean and sd (adjust if your format doesn't have a space) names = c("{col}_mean", "{col}_sd"), # Names for the split columns too_few = "align_start" # Handle any rows where formatting might be off ) %>% # Keep only the mean columns select(ends_with("_mean"))
Key Notes
- Column Selection: Adjust
starts_with("feature")to match your actual column names—usecontains("your_keyword"),matches("regex_pattern"), oreverything()if all columns except IDs are features. - Format Consistency: Make sure all feature columns follow the same
mean (sd)format. If some have no space before the parenthesis, change thedelimto just(. - Numeric Conversion: Both methods convert the extracted mean to numeric, which is essential for downstream analysis.
内容的提问来源于stack exchange,提问作者Louie Quicksell

