如何用R按Student分组统计Response列中yes的占比?
按学生分组计算Yes响应占比并匹配原数据行
需求说明
统计包含yes/no的Response列,按Student分组后计算每组中yes的占比,将占比添加到原数据的对应行中。
输入数据
| Student | Response |
|---|---|
| S1 | yes |
| S2 | yes |
| S1 | no |
| S5 | yes |
| S5 | yes |
| S7 | no |
| S8 | no |
期望输出
| Student | Response | percentage |
|---|---|---|
| S1 | yes | 50% |
| S2 | yes | 100% |
| S1 | no | 50% |
| S5 | yes | 100% |
| S5 | yes | 100% |
| S7 | no | 0% |
| S8 | no | 0% |
原代码问题分析
你提供的代码存在多处语法和逻辑错误:
- 管道符
%>%书写不完整(多处缺少%) - 使用
summarize后会将分组数据压缩为每组一行,无法直接关联原数据的多行记录 filter的位置错误,会过滤掉非yes的分组数据,导致无法计算全部分组的占比n()函数不需要传入参数,n(Response)写法错误
正确解决方案
方法一:分组内直接计算(推荐)
使用dplyr的mutate在分组内直接计算占比,无需拆分合并数据:
library(dplyr) library(scales) df <- tibble( Student = c("S1", "S2", "S1", "S5", "S5", "S7", "S8"), Response = c("yes", "yes", "no", "yes", "yes", "no", "no") ) result <- df %>% group_by(Student) %>% mutate( yes_count = sum(Response == "yes"), total_count = n(), percentage = label_percent(accuracy = 1)(yes_count / total_count) ) %>% ungroup() %>% select(Student, Response, percentage)
方法二:先分组统计再合并
先计算每个学生的占比,再通过left_join关联回原数据:
library(dplyr) library(scales) df <- tibble( Student = c("S1", "S2", "S1", "S5", "S5", "S7", "S8"), Response = c("yes", "yes", "no", "yes", "yes", "no", "no") ) # 计算分组占比 percent_stats <- df %>% group_by(Student) %>% summarize( yes_count = sum(Response == "yes"), total_count = n(), percentage = label_percent(accuracy = 1)(yes_count / total_count), .groups = "drop" ) # 关联回原数据 result <- df %>% left_join(percent_stats, by = "Student") %>% select(Student, Response, percentage)
两种方法都能得到你期望的输出结果,方法一更简洁高效。
内容的提问来源于stack exchange,提问作者k34l
相关产品推荐
相关产品推荐

