如何使用dplyr的mutate across结合ifelse批量处理以G开头的列
解决方案
你可以结合mutate() + across()批量处理所有以"G"开头的列,用ifelse()完成值的替换逻辑,直接把这段逻辑追加到原数据处理的管道末尾即可,具体代码如下:
library(stringr) library(dplyr) library(fastDummies) # 设置随机种子保证结果可复现 set.seed(123) score <- sample(1:100,20,replace=TRUE) df <- data.frame(score) df <- df %>% mutate(grp = cut(score, breaks = c(-Inf, seq(0, 100, by = 20), Inf)), grp = str_c("G", as.integer(droplevels(grp)), '_', str_replace(grp, '\\((\\d+),(\\d+)\\]', '\\1_\\2'))) %>% dummy_cols("grp", remove_selected_columns = TRUE) %>% rename_with(~ str_remove(.x, 'grp_'), starts_with('grp_')) %>% # 批量处理G开头的列 mutate(across(starts_with("G"), ~ ifelse(.x == 1, score, NA)))
关键逻辑解释
starts_with("G"):精准筛选所有列名以"G"开头的列,作为across()的处理目标~ ifelse(.x == 1, score, NA):.x代表当前正在处理的列,当该列值为1时,替换为对应行的score值;否则设为NA
进阶匹配(可选)
如果担心有其他含"G"的列被误处理,可以用正则表达式更精准地匹配列名格式(比如G1_0_20这种结构):
mutate(across(matches("^G\\d+_\\d+_\\d+$"), ~ ifelse(.x == 1, score, NA)))
内容的提问来源于stack exchange,提问作者Bruh
相关产品推荐
相关产品推荐

