使用group_by结合自定义函数计算HE升学占比遇问题求助
问题描述
编写了progressionby19函数用于计算升入高等教育(HE)的学生占比,规则如下:
- 排除
Student.Attended为"No"的学生 - 需对重复出现的
Student.ID去重
单独调用函数时结果正常,但结合group_by(Type)分组计算时出现错误;添加分组判断逻辑后,分组的升学占比结果仍与其他统计项不匹配。
原始代码
example <- tibble(Student.ID = c("#001","#002","#003","#004","#005"), Student.Attended = c("Yes", "Yes", "No", "Yes", "Yes"), entryby19 = c("Yes", "Yes", "Yes", "Yes", "No"), Type = c("Exhibition", "Exhibition", "Mentoring", "Mentoring", "Mentoring")) progressionby19 <- function(.data) { total <- .data %>% filter(Student.Attended == "Yes") %>% summarise(count = n_distinct(Student.ID)) progressed <- .data %>% filter(Student.Attended == "Yes") %>% filter(entryby19 == "Yes") %>% summarise(count = n_distinct(Student.ID)) (progressed/total)*100 } # 单独调用正常 progressionby19(example) # 分组调用出错 example %>% group_by(Type) %>% summarise(learner_count = n_distinct(Student.ID), progression_rate = progressionby19(.))
添加分组判断后的函数
progressionby19 <- function(.data) { if (dplyr::is_grouped_df(.data)) { return(dplyr::do(.data, progressionby19(.))) } total <- .data %>% filter(Student.Attended == "Yes") %>% summarise(count = n_distinct(Student.ID)) progressed <- .data %>% filter(Student.Attended == "Yes") %>% filter(entryby19 == "Yes") %>% summarise(count = n_distinct(Student.ID)) (progressed/total)*100 } # 单独调用正常 progressionby19(example) # 分组调用结果仍不符合预期 example %>% group_by(Type) %>% summarise(learner_count = n_distinct(Student.ID), progression_rate = progressionby19(.))
问题原因
- 原始函数返回的是数据框而非标量值,在
group_by后的summarise中调用时,无法将数据框直接作为列值返回,导致错误。 - 添加分组判断后使用
dplyr::do,返回的是分组的结果数据框,同样无法嵌入到summarise的单个列中,导致结果不匹配。
解决方案
方案1:修改函数返回标量值
将函数改为返回单个数值,适配summarise的调用场景:
progressionby19 <- function(.data) { # 先去重并筛选参与的学生 valid_students <- .data %>% filter(Student.Attended == "Yes") %>% distinct(Student.ID, .keep_all = TRUE) total <- nrow(valid_students) progressed <- valid_students %>% filter(entryby19 == "Yes") %>% nrow() if (total == 0) return(0) # 避免除以0 (progressed / total) * 100 } # 分组调用测试 example %>% group_by(Type) %>% summarise(learner_count = n_distinct(Student.ID), progression_rate = progressionby19(.))
运行结果:
# A tibble: 2 × 3 Type learner_count progression_rate <chr> <int> <dbl> 1 Exhibition 2 100 2 Mentoring 3 66.7
方案2:直接用dplyr原生语法计算(无需自定义函数)
如果不需要复用函数,可直接在分组后完成计算,逻辑更清晰:
example %>% group_by(Type) %>% # 先对每个分组去重有效学生 filter(Student.Attended == "Yes") %>% distinct(Student.ID, .keep_all = TRUE) %>% summarise( learner_count = n_distinct(Student.ID), progression_rate = (sum(entryby19 == "Yes") / n()) * 100 )
运行结果和方案1一致。
内容的提问来源于stack exchange,提问作者Sophie Elizabeth
相关产品推荐
相关产品推荐

