为何R语言foo2()函数未将同模式study分组为同一列表元素?
foo2()函数分组不符合预期的原因分析
问题背景
foo2()的设计目的是依据不同study的变量变化模式对数据分组:当所选变量中某一变量(如var1.var2)会随其他所选变量(如sample_id)的部分行变化时,具有相同模式的study应被归为同一组。
在输入数据中:
- study为Mecartty的行:var1.var2随sample_id的部分行变化,sample_id随study的部分行变化
- study为Larson的行:呈现与Mecartty完全相同的变量依赖模式
但调用foo2(Data, study, sample_id, var1.var2)后,预期应得到1个列表元素(两组study归为一组),实际却输出了2个独立的列表元素。
代码示例
library(tidyverse) library(rlang) Data <- read.table(h=T, text = "study sample_id var1.var2 effect Mecartty 1 L2V.L2L 1 Mecartty 1 L2G.L2L 2 Mecartty 2 L2R.L2V 3 Mecartty 2 L2R.L2G 4 Mecartty 2 L2V.L2G 5 Larson 1 L2R.L2G 6 Larson 1 L2R.L2L 7 Larson 1 L2G.L2L 8 Larson 2 L2R.L2G 9 Larson 2 L2R.L2L 10 Larson 2 L2G.L2L 11") foo2 <- function(dat, study_col, ...) { dot_cols <- ensyms(...) str_cols <- purrr::map_chr(dot_cols, rlang::as_string) dat %>% dplyr::select({{study_col}}, !!! dot_cols) %>% dplyr::group_by({{study_col}}) %>% dplyr::mutate(grp = across(all_of(str_cols), ~ { tmp <- n_distinct(.) case_when(tmp == 1 ~ 1, tmp == n() ~ 2, tmp >1 & tmp < n() ~ 3, TRUE ~ 4) }) %>% purrr::reduce(stringr::str_c, collapse="")) %>% dplyr::ungroup(.) %>% dplyr::group_split(grp, .keep = FALSE) } # 调用函数 foo2(Data, study, sample_id, var1.var2)
原因解析
问题核心出在grp分组键的生成逻辑上:
- 函数先按study分组,对每个study内部的目标变量计算
n_distinct(.)(变量不同值的数量),再通过case_when生成对应编码,最后拼接成grp字符串。 - 两组study的
grp计算结果不同:- Mecartty组:共5行。sample_id的不同值数量为2(满足
1<2<5,编码3);var1.var2的不同值数量为5(等于总行数,编码2),最终grp为"32"。 - Larson组:共6行。sample_id的不同值数量为2(满足
1<2<6,编码3);var1.var2的不同值数量为3(满足1<3<6,编码3),最终grp为"33"。
- Mecartty组:共5行。sample_id的不同值数量为2(满足
- 由于两个study的
grp字符串不一致,group_split会将它们拆分为两个独立的列表元素。
简言之,当前函数是基于每个study内部的绝对行数与变量不同值数量的关系生成分组键,而非判断变量间的依赖模式(比如var1.var2是否随sample_id的部分行变化),这才导致了不符合预期的分组结果。
内容的提问来源于stack exchange,提问作者Simon Harmel
相关产品推荐
相关产品推荐

