基于cluster分组长度创建二元变量flow.type的代码问题求助
问题排查与解决方案
问题描述
需要创建名为flow.type的二元变量,规则为:当cluster分组的行数为1时,赋值为"0";分组行数大于1时,赋值为"1"。尝试以下代码未成功,附数据样本:
# to determine flow type from the clustered groups, use binary version of if/else statement # flow.type 1 = 'event' # flow.type 0 = 'non-event' y <- y %>% group_by(cluster) %>% mutate(flow.type = case_when(length(cluster)>1 ~ "1", TRUE ~ "0")) %>% ungroup()
数据样本:
structure(list(Station = c("1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051017", "1051021", "1051021", "1051021", "1051021", "1051021"), Site.Name = c("Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Carroll's Crossing", "Laura River at Broken Dam Station", "Laura River at Broken Dam Station", "Laura River at Broken Dam Station", "Laura River at Broken Dam Station", "Laura River at Broken Dam Station"), Date.Time = c("20/10/2017 7:45", "24/10/2017 10:57", "27/12/2019 9:15", "16/01/2020 9:32", "15/04/2020 9:45", "12/05/2020 14:30", "17/06/2020 15:55", "11/09/2020 9:16", "12/01/2021 19:44", "13/01/2021 12:00", "27/01/2021 15:59", "27/01/2021 16:29", "27/01/2021 17:00", "19/02/2021 9:30", "17/01/2022 10:17", "27/12/2019 8:10", "31/12/2019 8:30", "21/01/2020 14:25", "21/01/2020 14:47", "14/05/2020 15:15"), Date = structure(c(17459, 17463, 18257, 18277, 18367, 18394, 18430, 18516, 18639, 18640, 18654, 18654, 18654, 18677, 19009, 18257, 18261, 18282, 18282, 18396), class = "Date"), Datedigit = c(17459745, 174631057, 18257915, 18277932, 18367945, 183941430, 184301555, 18516916, 186391944, 186401200, 186541559, 186541629, 186541700, 18677930, 190091017, 18257810, 18261830, 182821425, 182821447, 183961515), Sampling.Year = structure(c(3L, 3L, 5L, 5L, 5L, 5L, 5L, 6L, 6L, 6L, 6L, 6L, 6L, 6L, 7L, 5L, 5L, 5L, 5L, 5L ), .Label = c("2015-2016", "2016-2017", "2017-2018", "2018-2019", "2019-2020", "2020-2021", "2021-2022", "2022-2023"), class = "factor"), Season = c("Dry", "Dry", "Wet", "Wet", "Wet", "Dry", "Dry", "Dry", "Wet", "Wet", "Wet", "Wet", "Wet", "Wet", "Wet", "Wet", "Wet", "Wet", "Wet", "Dry"), cluster = c(56, 57, 58, 59, 60, 61, 62, 63, 64, 64, 65, 65, 65, 66, 67, 66, 67, 68, 68, 69)), row.names = c(NA, -20L), class = c("tbl_df", "tbl", "data.frame"))
问题原因
原代码中使用length(cluster)获取分组行数虽然逻辑上可行,但dplyr提供了更可靠的专用函数n()来返回当前分组的观测数量;同时原代码用case_when处理二元判断场景不够简洁,易因语法细节出错。
正确实现方法
方法一:分组+n()函数
通过group_by(cluster)分组后,用n()直接获取每组行数,再用if_else完成二元判断:
library(dplyr) y <- y %>% group_by(cluster) %>% mutate(flow.type = if_else(n() > 1, "1", "0")) %>% ungroup()
方法二:使用add_count(无需分组mutate)
先通过add_count生成每组的行数,再基于该列判断赋值,最后可选择删除中间列:
y <- y %>% add_count(cluster, name = "cluster_size") %>% mutate(flow.type = if_else(cluster_size > 1, "1", "0")) %>% select(-cluster_size) # 可选,删除临时生成的cluster_size列
验证结果
运行上述代码后,样本数据中:
- 仅1行的cluster(如56、57、58等)对应的
flow.type为"0" - 多行的cluster(如64、65、68等)对应的
flow.type为"1",完全符合需求。
内容的提问来源于stack exchange,提问作者CatN
相关产品推荐
相关产品推荐

