You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R按Student分组统计Response列中yes的占比?

按学生分组计算Yes响应占比并匹配原数据行

需求说明

统计包含yes/no的Response列,按Student分组后计算每组中yes的占比,将占比添加到原数据的对应行中。

输入数据

StudentResponse
S1yes
S2yes
S1no
S5yes
S5yes
S7no
S8no

期望输出

StudentResponsepercentage
S1yes50%
S2yes100%
S1no50%
S5yes100%
S5yes100%
S7no0%
S8no0%

原代码问题分析

你提供的代码存在多处语法和逻辑错误:

  • 管道符%>%书写不完整(多处缺少%)
  • 使用summarize后会将分组数据压缩为每组一行,无法直接关联原数据的多行记录
  • filter的位置错误,会过滤掉非yes的分组数据,导致无法计算全部分组的占比
  • n()函数不需要传入参数,n(Response)写法错误

正确解决方案

方法一:分组内直接计算(推荐)

使用dplyr的mutate在分组内直接计算占比,无需拆分合并数据:

library(dplyr)
library(scales)

df <- tibble(
  Student = c("S1", "S2", "S1", "S5", "S5", "S7", "S8"),
  Response = c("yes", "yes", "no", "yes", "yes", "no", "no")
)

result <- df %>%
  group_by(Student) %>%
  mutate(
    yes_count = sum(Response == "yes"),
    total_count = n(),
    percentage = label_percent(accuracy = 1)(yes_count / total_count)
  ) %>%
  ungroup() %>%
  select(Student, Response, percentage)

方法二:先分组统计再合并

先计算每个学生的占比,再通过left_join关联回原数据:

library(dplyr)
library(scales)

df <- tibble(
  Student = c("S1", "S2", "S1", "S5", "S5", "S7", "S8"),
  Response = c("yes", "yes", "no", "yes", "yes", "no", "no")
)

# 计算分组占比
percent_stats <- df %>%
  group_by(Student) %>%
  summarize(
    yes_count = sum(Response == "yes"),
    total_count = n(),
    percentage = label_percent(accuracy = 1)(yes_count / total_count),
    .groups = "drop"
  )

# 关联回原数据
result <- df %>%
  left_join(percent_stats, by = "Student") %>%
  select(Student, Response, percentage)

两种方法都能得到你期望的输出结果,方法一更简洁高效。

内容的提问来源于stack exchange,提问作者k34l

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 19:55:16