如何用tidyverse实现按LOCATION计算TOPIC-RESP组合的分布占比
计算各LOCATION下TOPIC与RESP组合的分布占比(Tidyverse简洁方案)
问题背景
我有一个包含LOCATION、TOPIC、RESP三个变量的数据集,希望计算每个LOCATION下,各TOPIC与RESP组合的分布占比。目前已通过循环实现需求,但想寻求更简洁的tidyverse解决方案,核心是在dplyr操作中获取当前组对应的LOCATION的所有Freq总和,以此计算占比。
玩具数据及初始处理代码
responses <- data.frame(LOCATION = c("LOC_A", "LOC_A", "LOC_A", "LOC_A", "LOC_A", "LOC_A", "LOC_A", "LOC_A", "LOC_B", "LOC_B", "LOC_B", "LOC_B", "LOC_B", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C"), TOPIC = c("Dogs", "Dogs", "Dogs", "Dogs", "Dogs", "Dogs", "Lizards", "Lizards", "Lizards", "Lizards", "Lizards", "Lizards", "Lizards", "Lizards", "Lizards", "Lizards", "Snakes", "Snakes", "Snakes", "Snakes", "Snakes", "Snakes", "Dogs", "Snakes", "Dogs", "Snakes", "Dogs", "Snakes", "Dogs", "Snakes"), RESP = c("Agree", "Disagree", "Agree", "Disagree", "Agree", "Disagree", "Agree", "Disagree", "Agree", "Disagree", "Agree", "Disagree", "Neither", "Agree", "Neither", "Agree", "Neither", "Agree", "Neither", "Agree", "Neither", "Agree", "Agree", "Neither", "Agree", "Neither", "Agree", "Disagree", "Disagree", "Neither")) # 获取各组合的频数 distribution <- responses %>% table() %>% as.data.frame() %>% dplyr::arrange(LOCATION, TOPIC, RESP)
已实现的循环解决方案
# 循环实现(不够简洁) out <- list() for(loc in unique(distribution$LOCATION)){ thisDist <- dplyr::filter(distribution, LOCATION == loc) thisDist$percent <- thisDist$Freq/sum(thisDist$Freq) out[[loc]] <- thisDist } out <- do.call("rbind", out)
Tidyverse简洁解决方案
最优方案:按LOCATION分组后直接计算占比
由于distribution中每一行对应唯一的LOCATION-TOPIC-RESP组合,只需按LOCATION分组,计算组内Freq的总和(即当前LOCATION的总样本数),再用每行的Freq除以该总和即可得到占比:
out <- distribution %>% dplyr::group_by(LOCATION) %>% dplyr::mutate(percent = Freq / sum(Freq)) %>% dplyr::ungroup() %>% dplyr::arrange(LOCATION, TOPIC, RESP)
备选方案:使用窗口函数明确指定分组(适用于更复杂场景)
如果需要在summarise中操作,可利用dplyr的窗口分组语法(需dplyr 1.1.0+),明确指定按LOCATION计算总频数:
out <- distribution %>% dplyr::group_by(LOCATION, TOPIC, RESP) %>% dplyr::summarise( Freq = dplyr::first(Freq), percent = Freq / sum(Freq, .by = LOCATION) ) %>% dplyr::ungroup() %>% dplyr::arrange(LOCATION, TOPIC, RESP)
内容的提问来源于stack exchange,提问作者Eliot Dixon
相关产品推荐
相关产品推荐

