如何在R中用dplyr的summarise获取因子变量的前后测合并计数
问题描述
我有一个以因子变量为主的数据集,想在R中使用dplyr包的summarise函数汇总其计数。数据集来自处理前后的场景,post组可能因响应缺失部分水平。
目前我能分别获取pre和post的计数,代码如下:
bla = data.frame(pre = c("a", "b", "c", "d", "e"), post = c("b", "d", "a", "a", "e")) bla$pre = as.factor(bla$pre) bla$post = as.factor(bla$post) bla %>% group_by(pre) %>% summarise(Count = n()) bla %>% group_by(post) %>% summarise(Count = n())
这会得到两个单独的计数表,但我希望得到如下格式的合并结果:
| Level | pre Count | post Count |
|---|---|---|
| a | 1 | 2 |
| b | 1 | 1 |
| c | 1 | 0 |
| d | 1 | 1 |
| e | 1 | 1 |
解决方案
以下两种方法都能帮你得到合并后的计数表,同时保留所有因子水平(post中缺失的水平会自动补0):
方法1:dplyr全连接补0
library(dplyr) # 统计pre组计数 pre_counts <- bla %>% group_by(Level = pre) %>% summarise(`pre Count` = n()) # 统计post组计数并补全所有pre水平 post_counts <- bla %>% group_by(Level = post) %>% summarise(`post Count` = n()) %>% right_join(tibble(Level = levels(bla$pre)), by = "Level") %>% mutate(`post Count` = replace_na(`post Count`, 0)) # 合并两个计数表并排序 final_table <- pre_counts %>% full_join(post_counts, by = "Level") %>% arrange(Level) print(final_table)
方法2:tidyr重塑数据后统计
这种方式通过先把宽格式转成长格式,再一次性统计所有水平的计数,代码更简洁:
library(dplyr) library(tidyr) final_table <- bla %>% # 将pre和post列转成长格式 pivot_longer(cols = c(pre, post), names_to = "Group", values_to = "Level") %>% # 按Group和Level统计计数 count(Group, Level, name = "Count") %>% # 转回宽格式,生成对应列名 pivot_wider(names_from = Group, values_from = Count, names_glue = "{Group} Count") %>% # 补全所有pre的因子水平 right_join(tibble(Level = levels(bla$pre)), by = "Level") %>% # 将NA替换为0 mutate(across(c(`pre Count`, `post Count`), ~replace_na(.x, 0))) %>% # 按Level排序 arrange(Level) print(final_table)
内容的提问来源于stack exchange,提问作者AirIcarus
相关产品推荐
相关产品推荐

