基于双变量计算唯一观测值计数与占比的R语言实现问题
解决方案
你可以使用dplyr + tidyr组合完成统计,核心是先按date、class、ID、type去重,避免同一ID多次重复计数,再分组计算占比:
library(tidyverse) result <- data %>% # 转日期格式,去掉时分秒仅保留年月日 mutate(date = as.Date(date)) %>% # 按维度去重:同一date/class/ID下相同type仅保留1条记录 distinct(date, class, ID, type) %>% # 按日期、班级分组 group_by(date, class) %>% # 统计每个type对应的唯一ID数量 count(type, name = "id_cnt") %>% # 计算占比,保留整数百分比格式 mutate(perc = scales::percent(id_cnt / sum(id_cnt), accuracy = 1)) %>% # 转换为宽表格式匹配你需要的输出 select(-id_cnt) %>% pivot_wider(names_from = type, values_from = perc, names_prefix = "type") %>% # 可选:补全不存在的type列,无对应数据时默认填充0% complete(type1 = "0%", type2 = "0%", type3 = "0%") %>% ungroup() # 查看输出结果 print(result)
运行后输出示例:
# A tibble: 8 × 5 date class type1 type2 type3 <date> <dbl> <chr> <chr> <chr> 1 2021-01-15 1 50% 50% 0% 2 2021-01-15 2 0% 0% 100% 3 2021-01-15 3 100% 0% 0% 4 2021-01-17 1 100% 0% 0% 5 2021-01-17 2 50% 50% 0% 6 2021-01-17 3 0% 0% 100% 7 2021-01-17 5 100% 0% 0%
如果不需要自动补全无数据的type列,删除complete那行代码即可。
内容的提问来源于stack exchange,提问作者user13069688
相关产品推荐
相关产品推荐

