如何基于table()生成的对象用dplyr添加占比列
基于table对象用dplyr添加占比列的实现方法
问题背景
用户拥有如下数据集:
dput(head(data, 50)) structure(list(Treatment = c("A", "A", "B", "A", "A", "A", "A", "B", "B", "B", "B", "A", "A", "B", "A", "B", "A", "B", "A", "A", "A", "A", "B", "B", "B", "A", "B", "B", "A", "B", "B", "B", "A", "A", "A", "A", "A", "A", "A", "A", "A", "A", "B", "A", "B", "A", "B", "A", "A", "B"), Death = c(1, 1, 1, 1, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 0, 0, 1, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0, 1, 1, 0, 1, 0, 1, 0, 0, 1, 0, 1, 1, 0, 1, 1, 1, 1, 1, 0, 0, 1, 1, 1), `Last observation (days)` = c(276, 212, 154, 222, 33, 299, 344, 180, 49, 324, 74, 66, 196, 269, 353, 332, 302, 211, 69, 55, 338, 103, 108, 7, 199, 64, 10, 236, 82, 242, 34, 239, 197, 315, 243, 5, 126, 44, 260, 363, 246, 193, 190, 151, 279, 215, 142, 183, 328, 119), `Age (years)` = c(92.64, 15.68, 10.39, 66.43, 79.59, 74.24, 77.06, 31.06, 11.28, 52.65, 16.66, 13.01, 42.91, 63.8, 9.99, 1.92, 33.52, 8.68, 61.97, 28.99, 86.73, 16.96, 5.8, 51.27, 21.28, 36.08, 26.12, 64.53, 52.99, 7.17, 42.37, 57.63, 83.48, 67.67, 1.12, 23.16, 81.61, 6.47, 72.69, 29.15, 73.69, 60.3, 9.21, 18.6, 34.73, 24.31, 0.37, 22.06, 9.89, 30.78), `Age (years)_cat` = c("old", "young", "young", "old", "old", "old", "old", "young", "young", "old", "young", "young", "young", "old", "young", "young", "young", "young", "old", "young", "old", "young", "young", "old", "young", "young", "young", "old", "old", "young", "young", "old", "old", "old", "young", "young", "old", "young", "old", "young", "old", "old", "young", "young", "young", "young", "young", "young", "young", "young")), row.names = c(NA, -50L), class = c("tbl_df", "tbl", "data.frame"))
已通过以下代码生成交叉表:
data %>% mutate(Death = ifelse(Death == 0, 'No Death', 'Death'))%$% as.data.frame(with(., table(Death, Treatment)))
目前使用以下代码实现添加占比列:
data %>% mutate(Death = ifelse(Death == 0, 'No Death', 'Death')) %>% group_by(Treatment, Death) %>% summarize(n = n()) %>% mutate(freq = n/sum(n))
问题:是否可以基于上述table()生成的对象,使用dplyr实现添加对应占比列的效果?
解决方案
可以直接基于table生成的数据框,通过dplyr完成占比列的添加,具体实现有两种方式:
方式一:分步处理
先将table对象转换为数据框,再分组计算占比:
# 生成table对应的数据集 table_df <- data %>% mutate(Death = ifelse(Death == 0, 'No Death', 'Death'))%$% as.data.frame(with(., table(Death, Treatment))) # 基于table_df添加占比列 table_df %>% group_by(Treatment) %>% mutate(freq = Freq / sum(Freq))
方式二:管道链式处理
将生成table数据框和计算占比合并为一个管道操作:
data %>% mutate(Death = ifelse(Death == 0, 'No Death', 'Death'))%$% as.data.frame(with(., table(Death, Treatment))) %>% group_by(Treatment) %>% mutate(freq = Freq / sum(Freq))
说明
这里核心逻辑是按Treatment分组,用每组的Freq总和(即该处理组的总样本数)去除每个Death类别的频数,得到的结果和你之前通过group_by(Treatment, Death)再summarize的结果完全一致。
内容的提问来源于stack exchange,提问作者12666727b9
相关产品推荐
相关产品推荐

