R语言中基于FactorCol1级别自动生成分组条件聚合列时值被错误覆盖的问题
嵌套map_dfc生成衍生列时的赋值错误分析
你遇到的问题核心是内部transmute里对同一个列名重复赋值,同时逻辑上混淆了FactorCol1级别和对应统计列的关联,导致最终列值被错误覆盖。
错误代码的问题拆解
看你内部嵌套的map_dfc代码:
map_dfc(levels(Data_Frame$FactorCol1), function(.y) Data_Frame %>% group_by(Col1) %>% transmute( !! sprintf("Last%dCol7%s", .x, .y) := sum(Col8[Col7 <= .x]), !! sprintf("Last%dCol7%s", .x, .y) := sum(Col9[Col7 <= .x]) ) %>% ungroup %>% select(-Col1) )
这里有两个关键错误:
- 重复赋值同一列名:在同一个
transmute调用里,你用完全相同的sprintf表达式生成了两次列名(比如当.x=1、.y="active"时,两次都是Last1Col7active),后一次赋值会直接覆盖前一次的结果,所以最终Last1Col7active会被设为sum(Col9[...])的值,也就是inactive的统计数,这和你的预期完全相反。 - 逻辑匹配错误:你遍历
FactorCol1的级别(active/inactive),但在生成列时,不管当前.y是哪个级别,都同时用了Col8(对应active)和Col9(对应inactive),这完全没必要——遍历级别应该是为了对应到各自的统计列,而不是重复生成同一对列。
修正后的代码
我们可以调整嵌套逻辑,让每个级别对应生成自己的统计列,避免重复赋值:
# 先定义要处理的阈值和级别 thresholds <- c(1, 2, 5, 10) status_levels <- levels(Data_Frame$FactorCol1) # 嵌套map生成正确的列 new_cols <- map_dfc(thresholds, function(.x) { map_dfc(status_levels, function(.y) { # 根据当前级别选择对应的Col8/Col9 target_col <- if (.y == "active") "Col8" else "Col9" Data_Frame %>% group_by(Col1) %>% transmute(!! sprintf("Last%dCol7%s", .x, .y) := sum(.data[[target_col]][Col7 <= .x])) %>% ungroup() %>% select(-Col1) }) }) # 合并到原数据框 Data_Frame <- bind_cols(Data_Frame, new_cols)
或者更简洁的方式,利用crossing生成所有阈值和级别的组合,再批量处理:
library(tidyverse) # 生成所有阈值和级别的组合 param_combinations <- crossing(threshold = c(1,2,5,10), status = levels(Data_Frame$FactorCol1)) # 批量生成列 new_cols <- param_combinations %>% pmap_dfc(function(threshold, status) { target_col <- if (status == "active") "Col8" else "Col9" Data_Frame %>% group_by(Col1) %>% transmute(!! glue::glue("Last{threshold}Col7{status}") := sum(.data[[target_col]][Col7 <= threshold])) %>% ungroup() %>% select(-Col1) }) Data_Frame <- bind_cols(Data_Frame, new_cols)
这样就能正确生成8个目标列,每个列对应各自的阈值和状态统计。
内容的提问来源于stack exchange,提问作者Ray
相关产品推荐
相关产品推荐

