基于被试Early条件试次量生成Late条件子集并计算均值的R实现求助
Hey there! Let's tackle this problem step by step using R. I'll break down each requirement and show you practical, readable code to make it work—including the dynamic subset sizing you're stuck on.
1. Count Trials per Condition per Subject
First, let's calculate how many Early and Late trials each subject has. We'll use dplyr for clean grouped operations:
library(dplyr) # First, let's formalize your sample data into a data frame df <- data.frame( subj = rep(1:3, each = 10), trial = rep(1:10, 3), factor = c( rep("Early", 4), rep("Late", 6), # Subject 1 rep("Early", 3), rep("Late", 7), # Subject 2 rep("Early", 2), rep("Late", 8) # Subject 3 ), rt = c(324,405,293,738,310,389,350,782,513,401, 420,230,309,456,241,400,189,329,519,230, 299,499,403,389,356,365,234,345,300,402) ) # Count trials per condition per subject trial_counts <- df %>% group_by(subj, factor) %>% summarize(trial_count = n(), .groups = "drop") %>% tidyr::pivot_wider(names_from = factor, values_from = trial_count) print(trial_counts)
This will output exactly the counts you specified:
# A tibble: 3 × 3 subj Early Late <int> <int> <int> 1 1 4 6 2 2 3 7 3 3 2 8
2. Calculate Mean of Late Trial Subset Means
The core challenge here is dynamically using each subject's Early trial count as the subset size for Late trials. expand.grid() isn't the right tool here (it creates Cartesian products, not unique trial combinations)—instead, we'll use combn(), which generates all possible unordered subsets of a given size from a vector.
We'll use dplyr::group_modify() to apply a custom function to each subject's data individually:
library(purrr) # For vectorized operations # Define a function to process one subject's data process_subject <- function(subj_data, subj_info) { # Get the number of Early trials (our subset size for Late) n_early <- subj_data %>% filter(factor == "Early") %>% nrow() # Extract all Late trial RTs for this subject late_rts <- subj_data %>% filter(factor == "Late") %>% pull(rt) # Generate all n-sized combinations of Late RTs, calculate each combination's mean subset_means <- combn(late_rts, n_early, FUN = mean) # Calculate the overall mean of these subset means overall_late_mean <- mean(subset_means) # Return results as a tidy tibble tibble( subj = unique(subj_data$subj), n_early = n_early, late_subset_overall_mean = overall_late_mean ) } # Apply the function to each subject late_subset_results <- df %>% group_by(subj) %>% group_modify(process_subject) %>% ungroup() print(late_subset_results)
Quick note on combn() vs expand.grid()
expand.grid() creates every possible ordered combination (allowing repeated trials and different sequences), which isn't what we want here. combn() generates only unique, unordered subsets of distinct trials—perfect for your requirement of sampling subsets matching the Early trial count without repetition.
3. Combine All Results (Optional)
If you want a single table with trial counts and subset mean results, merge the two data frames:
final_results <- trial_counts %>% left_join(late_subset_results, by = "subj") print(final_results)
内容的提问来源于stack exchange,提问作者arvag19

