You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中用dplyr的summarise获取因子变量的前后测合并计数

问题描述

我有一个以因子变量为主的数据集,想在R中使用dplyr包的summarise函数汇总其计数。数据集来自处理前后的场景,post组可能因响应缺失部分水平。

目前我能分别获取pre和post的计数,代码如下:

bla = data.frame(pre = c("a", "b", "c", "d", "e"), post = c("b", "d", "a", "a", "e"))
bla$pre = as.factor(bla$pre)
bla$post = as.factor(bla$post)
bla %>% group_by(pre) %>% summarise(Count = n())
bla %>% group_by(post) %>% summarise(Count = n())

这会得到两个单独的计数表,但我希望得到如下格式的合并结果:

Levelpre Countpost Count
a12
b11
c10
d11
e11
解决方案

以下两种方法都能帮你得到合并后的计数表,同时保留所有因子水平(post中缺失的水平会自动补0):

方法1:dplyr全连接补0

library(dplyr)

# 统计pre组计数
pre_counts <- bla %>% 
  group_by(Level = pre) %>% 
  summarise(`pre Count` = n())

# 统计post组计数并补全所有pre水平
post_counts <- bla %>% 
  group_by(Level = post) %>% 
  summarise(`post Count` = n()) %>%
  right_join(tibble(Level = levels(bla$pre)), by = "Level") %>%
  mutate(`post Count` = replace_na(`post Count`, 0))

# 合并两个计数表并排序
final_table <- pre_counts %>% 
  full_join(post_counts, by = "Level") %>%
  arrange(Level)

print(final_table)

方法2:tidyr重塑数据后统计

这种方式通过先把宽格式转成长格式,再一次性统计所有水平的计数,代码更简洁:

library(dplyr)
library(tidyr)

final_table <- bla %>%
  # 将pre和post列转成长格式
  pivot_longer(cols = c(pre, post), names_to = "Group", values_to = "Level") %>%
  # 按Group和Level统计计数
  count(Group, Level, name = "Count") %>%
  # 转回宽格式,生成对应列名
  pivot_wider(names_from = Group, values_from = Count, names_glue = "{Group} Count") %>%
  # 补全所有pre的因子水平
  right_join(tibble(Level = levels(bla$pre)), by = "Level") %>%
  # 将NA替换为0
  mutate(across(c(`pre Count`, `post Count`), ~replace_na(.x, 0))) %>%
  # 按Level排序
  arrange(Level)

print(final_table)

内容的提问来源于stack exchange,提问作者AirIcarus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 14:10:34