如何基于总均值±2SD筛选排除被试RT数据?(R语言)
Hey there! Let's work through this step by step to get your subject exclusion logic working correctly. First, I spotted a small misstep in your current code for calculating the grand mean and standard deviation—you’re using sd(sd_sub), which computes the standard deviation of subject-level RT standard deviations, but what we actually need is the standard deviation of the subject-level average RTs to set our exclusion threshold.
Here's a complete, corrected workflow using dplyr:
Step 1: Load required library and define your raw data
First, make sure dplyr is loaded, and structure your raw data (I’ve included your sample data here for completeness):
library(dplyr) # Raw sample data df <- tibble( Subj = rep(1:3, each = 10), Trial = rep(1:10, 3), factor = c("Early","Late","Early","Late","Early","Early","Late","Late","Late","Early", "Early","Late","Early","Late","Early","Early","Late","Late","Late","Early", "Early","Late","Early","Late","Early","Early","Late","Late","Late","Early"), rt = c(324,405,1093,738,310,389,1050,782,513,401, 920,1230,1509,1456,2341,900,1100,1329,1189,1030, 999,499,1003,389,356,365,1234,345,300,402) )
Step 2: Calculate subject-level statistics
Compute each subject’s average RT (and other useful stats) as you started:
subj_stats <- df %>% group_by(Subj) %>% summarise( N = n(), avgrt = mean(rt), sd_sub = sd(rt), min_rt = min(rt), median_rt = median(rt), max_rt = max(rt) ) %>% ungroup() # Ungroup to avoid issues with subsequent calculations
Step 3: Compute grand mean and threshold values
Calculate the overall mean of subject-level average RTs, plus the 2-standard-deviation threshold for exclusion:
# Get grand mean and grand standard deviation (of subject averages) grand_stats <- subj_stats %>% summarise( grand_mean = mean(avgrt), grand_sd = sd(avgrt) ) # Extract threshold values for filtering grand_mean <- grand_stats$grand_mean two_sd_range <- 2 * grand_stats$grand_sd
Step 4: Filter valid subjects
Identify which subjects have average RTs within the grand_mean ± 2*grand_sd range:
valid_subjs <- subj_stats %>% filter(avgrt >= grand_mean - two_sd_range, avgrt <= grand_mean + two_sd_range) # View which subjects are retained valid_subjs
Step 5: Get clean dataset with valid subjects only
Filter your original raw data to keep only the subjects that passed the exclusion criteria:
data_clean <- df %>% filter(Subj %in% valid_subjs$Subj) # Inspect the cleaned data data_clean
Key Notes
- The critical fix here is using
sd(avgrt)instead ofsd(sd_sub)—we need the spread of subject averages, not the spread of subject-level RT variability, to set our exclusion threshold. - This workflow ensures you exclude entire subjects whose overall average RT falls outside the 2SD range relative to the grand mean, regardless of individual trial or factor-level data.
内容的提问来源于stack exchange,提问作者arvag19

