如何用tidyverse将列转为多布尔列?求简化重复代码方案
问题描述
我有多组带时间后缀的category列,希望用mutate()和across()将其转换为每个类别对应的布尔列。目前通过重复3次相同代码实现了功能,但如果类别数量增多,这种方式会非常繁琐。尝试用map()简化代码但未成功,推测问题出在across()的.names参数无法与case_when()中使用的字符串相关联,尝试的简化代码无法运行,求助更简洁的实现方法。
原始实现代码
data <- data.frame(category_t1 = c("A","B","C","C","A","B"), category_t2 = c("A","C","B","B","B",NA), category_t3 = c("C","C",NA,"B",NA,"A")) data %>% mutate(across(starts_with("category"), ~case_when(.x == "A" ~ TRUE, !is.na(.x) ~ FALSE), .names = "{str_replace(.col, 'category', 'A')}"), across(starts_with("category"), ~case_when(.x == "B" ~ TRUE, !is.na(.x) ~ FALSE), .names = "{str_replace(.col, 'category', 'B')}"), across(starts_with("category"), ~case_when(.x == "C" ~ TRUE, !is.na(.x) ~ FALSE), .names = "{str_replace(.col, 'category', 'C')}"))
运行后输出
category_t1 category_t2 category_t3 A_t1 A_t2 A_t3 B_t1 B_t2 B_t3 C_t1 C_t2 1 A A C TRUE TRUE FALSE FALSE FALSE FALSE FALSE FALSE 2 B C C FALSE FALSE FALSE TRUE FALSE FALSE FALSE TRUE 3 C B <NA> FALSE FALSE NA FALSE TRUE NA TRUE FALSE 4 C B B FALSE FALSE FALSE FALSE TRUE TRUE TRUE FALSE 5 A B <NA> TRUE FALSE NA FALSE TRUE NA FALSE FALSE 6 B <NA> A FALSE NA TRUE TRUE NA FALSE FALSE NA
尝试的简化代码(无法运行)
data %>% mutate(across(starts_with("category"), ~case_when(.x == mod ~ TRUE, !is.na(.x) ~ FALSE), mod = levels(as.factor(data$category_t1)), .names = "{str_replace(.col, 'category', mod)}"))
解决方案
方法1:用purrr::map批量生成across调用
这是最贴近原始思路的简化方式,先提取所有唯一类别,再通过map遍历每个类别自动生成对应的across操作,避免重复代码:
library(dplyr) library(purrr) library(stringr) data <- data.frame(category_t1 = c("A","B","C","C","A","B"), category_t2 = c("A","C","B","B","B",NA), category_t3 = c("C","C",NA,"B",NA,"A")) # 提取所有非NA的唯一类别 unique_cats <- unique(unlist(data[, starts_with("category")])) %>% na.omit() # 遍历每个类别,生成对应的across逻辑 cat_across <- map(unique_cats, function(cat) { across(starts_with("category"), ~case_when(.x == cat ~ TRUE, !is.na(.x) ~ FALSE), .names = str_replace(.col, "category", cat)) }) # 展开所有across操作并执行mutate data %>% mutate(!!!cat_across)
方法2:长表转宽表(直观易读)
通过长格式转换统一处理类别匹配,再转回宽格式,逻辑更清晰:
library(dplyr) library(tidyr) data %>% # 将category列转成长表,保留时间后缀信息 pivot_longer(starts_with("category"), names_to = "time", values_to = "category") %>% # 为每个类别生成布尔标记 mutate( A = category == "A" & !is.na(category), B = category == "B" & !is.na(category), C = category == "C" & !is.na(category) ) %>% # 转回宽格式,还原原始列结构 pivot_wider( names_from = time, values_from = c(category, A, B, C), names_sep = "_" ) %>% # 调整列顺序与原始输出一致 select(starts_with("category"), starts_with("A"), starts_with("B"), starts_with("C"))
内容的提问来源于stack exchange,提问作者jrdavalos
相关产品推荐
相关产品推荐

