按ID重塑DataFrame:R语言分组统计需求实现求助
Your goal is to create a summary table grouped by id that counts age groups, education levels, and calculates the median blood value. The issue with your original code is that ddply + spread isn't the right tool for this task—instead, we can use dplyr's group_by() and summarize() functions to directly compute each required metric.
Step 1: Prepare the Sample Data
First, let's define your dataset for reproducibility:
dat <- data.frame( id = c(1, 1, 1, 2, 2, 2), age = c("30-39", "20-29", "30-39", "30-39", "20-29", "20-29"), edu = c("Primary", "Secondary", "Primary", "Primary", "Secondary", "Secondary"), blood = c(5.5, 8.7, 10, 11, 10, 9) )
Step 2: Correct dplyr Code
Use group_by() to group the data by id, then summarize() to compute each summary statistic:
library(dplyr) summary_df <- dat %>% group_by(id) %>% summarize( age30_39count = sum(age == "30-39"), # Count occurrences of "30-39" age group age20_29count = sum(age == "20-29"), # Count occurrences of "20-29" age group edu_pri_count = sum(edu == "Primary"), # Count Primary education entries edu_sec_count = sum(edu == "Secondary"),# Count Secondary education entries blood_median = median(blood) # Calculate median blood value ) %>% ungroup() # Optional: Remove grouping if not needed for further operations
Step 3: View the Result
Running this code will produce exactly your desired output:
print(summary_df) #> # A tibble: 2 × 6 #> id age30_39count age20_29count edu_pri_count edu_sec_count blood_median #> <dbl> <int> <int> <int> <int> <dbl> #> 1 1 2 1 2 1 8.7 #> 2 2 1 2 1 2 10
Why Your Original Code Didn't Work
The spread() function is designed to reshape long data into wide format by pivoting columns, but it doesn't handle counting or median calculations directly. By using summarize() with explicit count functions (sum(condition)), we can directly compute the metrics you need in a single, readable pipeline.
内容的提问来源于stack exchange,提问作者Rudro88

