You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量处理数据框,提取所有特征列的mean并移除括号内数据?

Batch Extract Mean Values from Multiple Feature Columns

Got it, let's solve this efficiently—no need to repeat code for 109 columns! Your goal is to pull out the mean values from every feature column (each containing both mean and sd, formatted like mean_value (sd_value)), and we can do this in bulk using tidyverse tools in R.

Step 1: Load Required Packages

First, make sure you have the tidyverse installed and loaded—it combines dplyr for data manipulation and tidyr for reshaping data, which is perfect for this task:

library(tidyverse)

Step 2: Batch Extract Mean Values (Two Simple Methods)

Method 1: Using across() + str_extract()

This method uses regex to directly pull the mean value (everything before the first left parenthesis) from every target column, then converts it to a numeric value.

Suppose your feature columns have a consistent prefix like feature_ (adjust the selector if your columns are named differently—use everything() if all non-ID columns are features):

# Extract mean values and create new columns with "_mean" suffix
df_mean_only <- df %>%
  mutate(
    across(
      .cols = starts_with("feature"), # Select all feature columns (adjust this to match your column names)
      .fns = ~ str_extract(., "^[^(]+") %>% trimws() %>% as.numeric(),
      .names = "{col}_mean"
    )
  ) %>%
  # Optional: Drop the original feature columns if you don't need them
  select(-starts_with("feature"))
  • ^[^(]+ is the regex that grabs all characters from the start of the string up to (but not including) the first (.
  • trimws() removes any extra spaces around the mean value (in case your formatting has gaps).
  • .names = "{col}_mean" ensures each new column keeps the original feature name plus a _mean suffix.

Method 2: Using separate_wider_delim() (More Intuitive for Splitting)

If you want to temporarily split columns into both mean and sd (then just keep the mean), separate_wider_delim is a cleaner option for bulk splitting:

df_mean_only <- df %>%
  separate_wider_delim(
    cols = starts_with("feature"), # Target feature columns
    delim = " (", # The separator between mean and sd (adjust if your format doesn't have a space)
    names = c("{col}_mean", "{col}_sd"), # Names for the split columns
    too_few = "align_start" # Handle any rows where formatting might be off
  ) %>%
  # Keep only the mean columns
  select(ends_with("_mean"))

Key Notes

  • Column Selection: Adjust starts_with("feature") to match your actual column names—use contains("your_keyword"), matches("regex_pattern"), or everything() if all columns except IDs are features.
  • Format Consistency: Make sure all feature columns follow the same mean (sd) format. If some have no space before the parenthesis, change the delim to just (.
  • Numeric Conversion: Both methods convert the extracted mean to numeric, which is essential for downstream analysis.

内容的提问来源于stack exchange,提问作者Louie Quicksell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:56:02