You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中按id_study分组统计edu列的NA值数量?

Hey there! Let's work through this problem together—you're dealing with a pretty large dataset (28M rows, 1258 unique studies), so we need efficient, reliable ways to identify the 728 groups where edu is entirely NA, plus count those NA values per group.

First, let's break down why your original aggregate_2 code didn't work: when you use FUN=length(dat[!is.na(dat)]), you're passing a precomputed value (not a function) to aggregate. The FUN argument needs a function that can run per group, not a single value calculated on the entire dataset.

Here are three solid solutions, ranging from base R to tools optimized for big data:

1. Base R (No extra packages needed)

We'll calculate total observations and non-NA counts per group, then filter for groups with zero non-NA values:

# Calculate total observations and non-NA `edu` counts per study
group_metrics <- aggregate(
  list(total_obs = dat$edu, edu_non_na = dat$edu),
  by = list(id_study = dat$id_study),
  FUN = function(x) c(total = length(x), non_na = sum(!is.na(x)))
)

# Convert the aggregated matrix columns into a clean data frame
group_metrics <- do.call(data.frame, group_metrics)

# Filter groups where `edu` is entirely NA (non-NA count = 0)
all_na_studies <- subset(group_metrics, edu_non_na.non_na == 0)

# The NA count per group is just the total observations for that group
all_na_studies$edu_na_count <- all_na_studies$total_obs.total

2. dplyr (Readable, efficient for large data)

If you prefer a more intuitive syntax, dplyr makes grouping and summarizing straightforward:

library(dplyr)

# Generate summary stats per study
study_summary <- dat %>%
  group_by(id_study) %>%
  summarise(
    total_obs = n(),
    edu_non_na = sum(!is.na(edu)),
    edu_na_count = sum(is.na(edu))
  ) %>%
  ungroup()

# Filter to get only studies with all NA `edu` values
all_na_studies <- study_summary %>%
  filter(edu_non_na == 0)

3. data.table (Fastest for huge datasets like yours)

For 28 million rows, data.table is the most efficient option—it’s optimized for speed and memory:

library(data.table)

# Convert your data frame to a data.table (in-place, no extra memory)
setDT(dat)

# Group and calculate stats in one step
study_summary <- dat[, .(
  total_obs = .N,
  edu_non_na = sum(!is.na(edu)),
  edu_na_count = sum(is.na(edu))
), by = id_study]

# Filter for all-NA `edu` groups
all_na_studies <- study_summary[edu_non_na == 0]

All three methods should give you exactly the 728 groups you're looking for, plus the count of NA values (which will equal the total observations for each group, since every row in those groups has an NA in edu).

内容的提问来源于stack exchange,提问作者Shivy b

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:16:13