You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr实现条件求和:统计特定年龄段未入学人数

问题:统计符合年龄条件的未入学家庭成员数量

我有一份家庭花名册数据集,需要统计每个家庭中年龄在6至17岁之间且edu_enrolment字段值为'no'的成员数量。目前我的代码会统计所有edu_enrolment为'no'的情况,没有考虑年龄条件,请求修正逻辑。

数据集

df <- structure(list(id = c(1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 4, 
4, 5, 5, 5, 5, 5, 5), age = c(47, 15, 8, 35, 30, 17, 5, 3, 23, 
15, 12, 4, 18, 10, 56, 41, 15, 12, 4, 3), edu_enrolment = c(NA, 
"no", "yes", NA, NA, "dnk", "yes", "yes", NA, "no", "no", "yes", 
NA, "yes", NA, NA, "yes", "no", "yes", "yes")), class = c("tbl_df", 
"tbl", "data.frame"), row.names = c(NA, -20L))

现有代码

education <- education |>
            group_by(id) |>
            mutate(
              notenroll_count = ifelse(all(is.na(edu_enrolment) | edu_enrolment %in% c('dnk')), NA, sum(edu_enrolment == 'no', na.rm = TRUE))
            ) |>
            ungroup()

修正后的代码

# 对齐数据集变量名,避免报错
education <- df |>
  group_by(id) |>
  mutate(
    notenroll_count = ifelse(
      all(is.na(edu_enrolment) | edu_enrolment %in% c('dnk')), 
      NA, 
      # 新增年龄范围判断,仅统计6-17岁且edu_enrolment为'no'的成员
      sum(age >= 6 & age <= 17 & edu_enrolment == 'no', na.rm = TRUE)
    )
  ) |>
  ungroup()

修正说明

  • 在统计条件中新增了age >= 6 & age <= 17的判断,确保只统计年龄在6至17岁之间的未入学成员
  • 补充了数据集变量名的对齐操作,原代码使用education但提供的数据集是df,避免运行时出现未定义变量的错误

内容的提问来源于stack exchange,提问作者Stephen Okiya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 23:21:17