You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用group_by结合自定义函数计算HE升学占比遇问题求助

问题描述

编写了progressionby19函数用于计算升入高等教育(HE)的学生占比,规则如下:

  • 排除Student.Attended为"No"的学生
  • 需对重复出现的Student.ID去重
    单独调用函数时结果正常,但结合group_by(Type)分组计算时出现错误;添加分组判断逻辑后,分组的升学占比结果仍与其他统计项不匹配。

原始代码

example <- tibble(Student.ID = c("#001","#002","#003","#004","#005"),
               Student.Attended = c("Yes", "Yes", "No", "Yes", "Yes"), 
               entryby19 = c("Yes", "Yes", "Yes", "Yes", "No"),
               Type = c("Exhibition", "Exhibition", "Mentoring", "Mentoring", "Mentoring"))

progressionby19 <- function(.data) {
  total <- .data %>%  
    filter(Student.Attended == "Yes") %>% 
    summarise(count = n_distinct(Student.ID))
  
  progressed <- .data %>%  
    filter(Student.Attended == "Yes") %>% 
    filter(entryby19 == "Yes") %>% 
    summarise(count = n_distinct(Student.ID))
  
  (progressed/total)*100
} 

# 单独调用正常
progressionby19(example)

# 分组调用出错
example %>% 
  group_by(Type) %>% 
  summarise(learner_count = n_distinct(Student.ID),
            progression_rate = progressionby19(.)) 

添加分组判断后的函数

progressionby19 <- function(.data) {
  
  if (dplyr::is_grouped_df(.data)) {
    return(dplyr::do(.data, progressionby19(.)))
  }
  
  total <- .data %>%  
    filter(Student.Attended == "Yes") %>% 
    summarise(count = n_distinct(Student.ID))
  
  progressed <- .data %>%  
    filter(Student.Attended == "Yes") %>% 
    filter(entryby19 == "Yes") %>% 
    summarise(count = n_distinct(Student.ID))
  
  (progressed/total)*100
}

# 单独调用正常
progressionby19(example)

# 分组调用结果仍不符合预期
example %>% 
  group_by(Type) %>% 
  summarise(learner_count = n_distinct(Student.ID),
            progression_rate = progressionby19(.)) 

问题原因

  1. 原始函数返回的是数据框而非标量值,在group_by后的summarise中调用时,无法将数据框直接作为列值返回,导致错误。
  2. 添加分组判断后使用dplyr::do,返回的是分组的结果数据框,同样无法嵌入到summarise的单个列中,导致结果不匹配。

解决方案

方案1:修改函数返回标量值

将函数改为返回单个数值,适配summarise的调用场景:

progressionby19 <- function(.data) {
  # 先去重并筛选参与的学生
  valid_students <- .data %>%
    filter(Student.Attended == "Yes") %>%
    distinct(Student.ID, .keep_all = TRUE)
  
  total <- nrow(valid_students)
  progressed <- valid_students %>%
    filter(entryby19 == "Yes") %>%
    nrow()
  
  if (total == 0) return(0) # 避免除以0
  (progressed / total) * 100
}

# 分组调用测试
example %>% 
  group_by(Type) %>% 
  summarise(learner_count = n_distinct(Student.ID),
            progression_rate = progressionby19(.)) 

运行结果:

# A tibble: 2 × 3
  Type        learner_count progression_rate
  <chr>               <int>            <dbl>
1 Exhibition              2              100
2 Mentoring               3             66.7

方案2:直接用dplyr原生语法计算(无需自定义函数)

如果不需要复用函数,可直接在分组后完成计算,逻辑更清晰:

example %>%
  group_by(Type) %>%
  # 先对每个分组去重有效学生
  filter(Student.Attended == "Yes") %>%
  distinct(Student.ID, .keep_all = TRUE) %>%
  summarise(
    learner_count = n_distinct(Student.ID),
    progression_rate = (sum(entryby19 == "Yes") / n()) * 100
  )

运行结果和方案1一致。

内容的提问来源于stack exchange,提问作者Sophie Elizabeth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 09:55:56