在tidyverse中基于现有列批量生成新列的高效方法
问题描述
现有如下数据集:
Squat1Kg Squat2Kg Squat3Kg Bench1Kg Bench2Kg Bench3Kg Deadlift1Kg Deadlift2Kg Deadlift3Kg <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> 1 75 80 -90 50 55 60 95 105 108. 2 95 100 105 62.5 67.5 -72.5 100 110 -120 3 85 90 100 55 62.5 -65 90 100 105 4 125 132 138. 115 122. -128. 150 165 170 5 80 85 90 40 50 -60 112. 120 125 6 90 -95 100 60 -65 -67.5 90 105 115 7 85 95 100 40 47.5 -50 115 130 140 8 210 225 232. 150 160 -165 240 260 -270
需要生成两类新列:
- 以
WeightTried_为前缀的新列,存储对应原列的绝对值,列名示例:[1] "WeightTried_Squat1Kg" "WeightTried_Squat2Kg" "WeightTried_Squat3Kg" [4] "WeightTried_Bench1Kg" "WeightTried_Bench2Kg" "WeightTried_Bench3Kg" [7] "WeightTried_Deadlift1Kg" "WeightTried_Deadlift2Kg" "WeightTried_Deadlift3Kg" - 以
Lifted为前缀、?为后缀的新列,标记原列值是否为正(正为1,否则为0),列名示例:LiftedSquat1Kg?、LiftedBench1Kg?等。
若使用mutate逐一编写过于繁琐,如何在tidyverse中高效实现上述需求?
解决方案
可以用dplyr的across()函数批量处理目标列,一次性生成两类新列,无需逐一编写重复代码:
library(tidyverse) # 假设原数据集名为df df_processed <- df %>% mutate( # 生成带WeightTried_前缀的绝对值列 across( everything(), # 处理所有列,若需指定列可替换为matches("Squat|Bench|Deadlift")等选择器 ~ abs(.x), .names = "WeightTried_{col}" ), # 生成带Lifted前缀、?后缀的标记列 across( everything(), ~ as.integer(.x > 0), # 正数值返回1,否则返回0 .names = "Lifted{col}?" ) )
关键说明:
across()函数是tidyverse中批量处理列的核心工具,能对指定列批量应用同一函数.names参数支持模板化生成新列名,{col}会自动替换为原列名,完美匹配需求的命名规则- 如果只需要处理特定列,把
everything()替换成列选择器即可,比如starts_with(c("Squat", "Bench", "Deadlift"))
内容的提问来源于stack exchange,提问作者Norhther
相关产品推荐
相关产品推荐

