如何轻松创建变量统计分组DataFrame中未接受处理d的行数?
Hey there! I get it—dealing with clunky group-wise calculations can be a real hassle. Let’s dive into a few concise Base R approaches to get n0, the count of cases not receiving treatment d for each id group in your DataFrame df1.
First, let’s set up sample data to mirror your scenario:
# Example data: id groups, where d=1 = treated, d=0 = untreated df1 <- data.frame( id = rep(c(1, 2, 3), each = 5), d = sample(c(0, 1), 15, replace = TRUE) )
Method 1: Use aggregate() (most intuitive for DataFrames)
This is Base R’s go-to function for group-by operations on DataFrames. We directly specify the column to analyze, the grouping variable, and a simple function to count untreated cases:
# Count 0s (untreated) per id group n0 <- aggregate(d ~ id, df1, function(x) sum(x == 0)) # Rename the result column for clarity colnames(n0)[2] <- "n0"
If your d column uses text labels (like "treated"/"untreated" instead of 0/1), just adjust the condition inside sum():
n0 <- aggregate(d ~ id, df1, function(x) sum(x == "untreated"))
Method 2: Use tapply() (super concise for vectors)
If you prefer working directly with vectors, tapply() is perfect. It returns a named vector by default, which you can easily convert to a tidy DataFrame:
# Get counts as a named vector n0_vec <- tapply(df1$d, df1$id, function(x) sum(x == 0)) # Convert to DataFrame n0 <- as.data.frame.table(n0_vec, responseName = "n0") colnames(n0)[1] <- "id"
Method 3: Use by() (flexible for group-wise operations)
by() splits your data into groups and applies a function to each. The result can be converted to a DataFrame with one quick line:
# Calculate counts per group n0_list <- by(df1$d, df1$id, function(x) sum(x == 0)) # Convert to DataFrame n0 <- data.frame(id = names(n0_list), n0 = as.vector(n0_list))
All these methods skip the hassle of manual loops or multi-step processing—they’re clean, readable, and fully Base R, so no extra packages required!
内容的提问来源于stack exchange,提问作者jay.sf

