You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用tidyverse实现按LOCATION计算TOPIC-RESP组合的分布占比

计算各LOCATION下TOPIC与RESP组合的分布占比(Tidyverse简洁方案)

问题背景

我有一个包含LOCATION、TOPIC、RESP三个变量的数据集,希望计算每个LOCATION下,各TOPIC与RESP组合的分布占比。目前已通过循环实现需求,但想寻求更简洁的tidyverse解决方案,核心是在dplyr操作中获取当前组对应的LOCATION的所有Freq总和,以此计算占比。


玩具数据及初始处理代码

responses <- data.frame(LOCATION = c("LOC_A", "LOC_A", "LOC_A", "LOC_A", "LOC_A", 
                                     "LOC_A", "LOC_A", "LOC_A", 
                                     "LOC_B", "LOC_B", "LOC_B", "LOC_B", "LOC_B", 
                                     "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", 
                                     "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", 
                                     "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C"),
                        TOPIC = c("Dogs", "Dogs", "Dogs", "Dogs", "Dogs", "Dogs", 
                                  "Lizards", "Lizards", "Lizards",
                                  "Lizards", "Lizards", "Lizards", "Lizards", "Lizards", 
                                   "Lizards", "Lizards", "Snakes", "Snakes", "Snakes", "Snakes", "Snakes", 
                                  "Snakes", "Dogs", "Snakes", "Dogs", "Snakes", "Dogs", 
                                  "Snakes", "Dogs", "Snakes"),
                        RESP = c("Agree", "Disagree", "Agree", "Disagree", "Agree", 
                                 "Disagree", "Agree", "Disagree", 
                                 "Agree", "Disagree", "Agree", "Disagree", "Neither", "Agree",
                                 "Neither", "Agree", "Neither", "Agree", "Neither", 
                                 "Agree", "Neither", "Agree", "Agree", "Neither", 
                                 "Agree", "Neither", "Agree", "Disagree", "Disagree",
                                 "Neither"))

# 获取各组合的频数
distribution <- responses %>% 
  table() %>% 
  as.data.frame() %>% 
  dplyr::arrange(LOCATION, TOPIC, RESP) 

已实现的循环解决方案

# 循环实现(不够简洁)
out <- list()
for(loc in unique(distribution$LOCATION)){
  thisDist <- dplyr::filter(distribution, LOCATION == loc)
  thisDist$percent <- thisDist$Freq/sum(thisDist$Freq)
  out[[loc]] <- thisDist
}
out <- do.call("rbind", out)

Tidyverse简洁解决方案

最优方案:按LOCATION分组后直接计算占比

由于distribution中每一行对应唯一的LOCATION-TOPIC-RESP组合,只需按LOCATION分组,计算组内Freq的总和(即当前LOCATION的总样本数),再用每行的Freq除以该总和即可得到占比:

out <- distribution %>%
  dplyr::group_by(LOCATION) %>%
  dplyr::mutate(percent = Freq / sum(Freq)) %>%
  dplyr::ungroup() %>%
  dplyr::arrange(LOCATION, TOPIC, RESP)

备选方案:使用窗口函数明确指定分组(适用于更复杂场景)

如果需要在summarise中操作,可利用dplyr的窗口分组语法(需dplyr 1.1.0+),明确指定按LOCATION计算总频数:

out <- distribution %>%
  dplyr::group_by(LOCATION, TOPIC, RESP) %>%
  dplyr::summarise(
    Freq = dplyr::first(Freq),
    percent = Freq / sum(Freq, .by = LOCATION)
  ) %>%
  dplyr::ungroup() %>%
  dplyr::arrange(LOCATION, TOPIC, RESP)

内容的提问来源于stack exchange,提问作者Eliot Dixon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 22:48:10