You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于分组变量值对数据框多列批量应用条件函数?

批量处理数值列生成新列(tidyverse方案)

你可以用dplyr中的across()函数实现批量处理,无需逐个手动定义新列。以下是两种简洁的实现方式:

方法一:使用if_else(适用于二元条件)

因为你的分组只有"a"和"b"两种情况,用if_else比case_when更简洁:

library(tidyverse)

df = tribble(
  ~x,     ~y,     ~z,
  1,     "a",     5,   
  2,     "b",     6,   
  3,     "a",     7,    
  4,     "b",     8,  
  5,     "a",     9,  
  6,     "b",     10
)

# 批量处理所有数值列,生成带_new后缀的新列
df_processed = df %>%
  mutate(across(where(is.numeric), 
                ~ .x + if_else(y == "a", 1, 2),
                .names = "{.col}_new"))

print(df_processed)

输出结果:

# A tibble: 6 × 5
      x y         z x_new z_new
  <dbl> <chr> <dbl> <dbl> <dbl>
1     1 a         5     2     6
2     2 b         6     4     8
3     3 a         7     4     8
4     4 b         8     6    10
5     5 a         9     6    10
6     6 b        10     8    12

方法二:使用case_when(适用于多条件扩展)

如果后续分组条件增多,用case_when更易扩展:

df_processed = df %>%
  mutate(across(where(is.numeric), 
                ~ case_when(
                  y == "a" ~ .x + 1,
                  y == "b" ~ .x + 2
                  # 可添加更多条件
                ),
                .names = "{.col}_new"))

自定义处理列

如果不需要处理所有数值列,可指定特定列(比如只处理x和z):

df_processed = df %>%
  mutate(across(c(x, z), 
                ~ .x + if_else(y == "a", 1, 2),
                .names = "{.col}_new"))

代码说明

  • where(is.numeric):自动筛选所有数值型列,无需手动列名
  • .x:代表across当前迭代处理的列
  • .names = "{.col}_new":自动生成新列名,原列名后加_new后缀

内容的提问来源于stack exchange,提问作者Meisam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 10:55:28