在R语言中实现同类对比表格转换技术求助
R语言表格转换解决方案
先把你的原始数据转换成R可处理的数据框:
library(tidyverse) # 原始数据 df <- tibble( ID = c(1,1,2,2,3,3,4,4,5,5,6,6,7,7,8,9), Condition = c("A","B","A","B","A","B","A","B","A","B","A","B","A","B","A","B"), Count = c(1,0,1,1,0,1,1,1,1,1,1,0,0,1,0,0) )
接下来分两步完成转换:
- 将长格式数据转为宽格式:用
pivot_wider把Condition的A、B转为列,缺失的对应值填充为0(比如ID8只有A的记录,ID9只有B的记录,转宽后对应列补0) - 统计每种A/B组合的ID数量:按A和B分组,统计每组的ID个数
完整代码:
result <- df %>% # 转宽格式,填充缺失值为0 pivot_wider( id_cols = ID, names_from = Condition, values_from = Count, values_fill = 0 ) %>% # 按A、B分组统计ID数量 count(A, B, name = "Count of ID") print(result)
运行后得到的结果就是你需要的表格:
# A tibble: 4 × 3 A B `Count of ID` <dbl> <dbl> <int> 1 0 0 1 2 0 1 2 3 1 0 2 4 1 1 3
如果不用tidyverse,也可以用基础R的方法:
# 转宽格式 wide_df <- reshape(df, idvar = "ID", timevar = "Condition", direction = "wide") # 填充缺失值为0 wide_df[is.na(wide_df)] <- 0 # 重命名列名 colnames(wide_df) <- c("ID", "A", "B") # 统计组合数量 result_base <- aggregate(ID ~ A + B, data = wide_df, FUN = length) colnames(result_base)[3] <- "Count of ID" print(result_base)
两种方法都能得到你想要的结果,tidyverse的方式代码更简洁易读。
内容的提问来源于stack exchange,提问作者Philip
相关产品推荐
相关产品推荐

