使用map+mutate处理tibble时行数异常倍增的问题排查
问题排查:行数变为3倍的原因
你使用purrr::map_dfr的核心问题是:这个函数的作用是行绑定(row-bind)每次函数调用的输出结果。你传入了3个类型参数(sofa/couch/settee),每个调用都会返回完整的原数据框+对应新列,三次行绑定后就会把原数据重复3次,最终行数变成原数据的3倍。而你实际需要的是在同一个数据框上逐步添加新列,而非堆叠多个数据框。
解决方案1:基于现有函数的修正调用
保留你写的cleaning_fcn,将map_dfr替换为purrr::reduce——它会把初始数据框依次传入每个函数调用,累积添加新列:
library(tidyverse) # 保留你的自定义函数 cleaning_fcn <- function(.df, .x){ .df %>% mutate(!!sym(paste0("explain_", .x)) := case_when( !!sym(paste0("sit_comfy_", .x ,"_1")) == 1 ~ "Just better", !!sym(paste0("sit_comfy_", .x, "_2")) == 1 ~ "Nice shape", !!sym(paste0("sit_comfy_", .x ,"_3")) == 1 ~ "Like the color", !!sym(paste0("sit_comfy_", .x ,"_4")) == 1 ~ "Nice material"), !!sym(paste0("explain_", .x)) := factor(!!sym(paste0("explain_", .x)), levels = c("Just better", "Nice shape", "Like the color", "Nice material"))) } # 用reduce实现累积添加列 result <- reduce(c("sofa", "couch", "settee"), cleaning_fcn, .init = d)
解决方案2:更简洁的tidyverse批量处理写法
无需自定义循环函数,通过数据重塑(pivot)批量处理所有类型,代码更简洁且易维护:
library(tidyverse) result <- d %>% # 将宽格式数据转成长格式,拆分列名获取类型和编号 pivot_longer( cols = everything(), names_to = c(NA, NA, "type", "num"), # 忽略前两个前缀部分 names_sep = "_", values_to = "value" ) %>% filter(value == 1) %>% # 只保留值为1的行(对应选中的选项) # 生成解释文本并转为指定顺序的因子 mutate( explain = case_match( num, "1" ~ "Just better", "2" ~ "Nice shape", "3" ~ "Like the color", "4" ~ "Nice material" ), explain = factor(explain, levels = c("Just better", "Nice shape", "Like the color", "Nice material")) ) %>% # 转回宽格式,生成explain_xxx列 pivot_wider( id_cols = row_number(), names_from = type, values_from = explain, names_prefix = "explain_" ) %>% select(-row_number()) %>% # 和原数据合并 bind_cols(d, .)
验证结果:result的行数与原数据d一致,且包含explain_sofa、explain_couch、explain_settee三个因子列,功能与你手动重复mutate的写法完全一致。
内容的提问来源于stack exchange,提问作者C.Robin
相关产品推荐
相关产品推荐

