如何在R中分别针对被试与目标单独剔除异常值并替换为NA
Hey there! Let's break this down clearly since you're new to R. First, I noticed your current tapply code only calculates means/SDs by either target or ptp alone—but you need to group by both ptp (subject) and target together. That's the key detail here!
Let's start by recreating your dataset so you can test the code directly:
# Create your sample data data <- data.frame( ptp = c(1,1,2,2,3,3,1,1,2,2,3,3), rt = c(1094,2100,918,1211,1142,1222,10,993,897,1078,4002,899), group = c("E","E","E","E","E","E","W","W","W","W","W","W"), target = c("dog","hat","dog","hat","dog","hat","dog","hat","dog","hat","dog","hat") )
Method 1: Base R (no extra packages needed)
We'll use the ave() function, which lets us calculate summary stats (mean/SD) for each group and returns a vector the same length as your original data—perfect for checking each rt value against its group's range.
# Calculate group-specific mean and SD (grouped by ptp + target) data$rt_mean <- ave(data$rt, list(data$ptp, data$target), FUN = mean) data$rt_sd <- ave(data$rt, list(data$ptp, data$target), FUN = sd) # Define the outlier range: mean ± 1.96*SD, replace outliers with NA data$rt_clean <- ifelse( data$rt >= (data$rt_mean - 1.96*data$rt_sd) & data$rt <= (data$rt_mean + 1.96*data$rt_sd), data$rt, NA ) # View the result data
Method 2: Tidyverse (dplyr package, more readable for beginners)
If you're open to using the dplyr package (part of the tidyverse, super popular for data manipulation), this approach is more intuitive for grouping operations:
First install/load the package if you haven't already:
install.packages("dplyr") library(dplyr)
Then clean your data with grouped operations:
data_clean <- data %>% group_by(ptp, target) %>% # Group by subject AND target mutate( rt_mean = mean(rt), rt_sd = sd(rt), rt_clean = ifelse( rt >= (rt_mean - 1.96*rt_sd) & rt <= (rt_mean + 1.96*rt_sd), rt, NA ) ) %>% ungroup() # Ungroup when done # View the cleaned data data_clean
What this does
Both methods will:
- For every combination of
ptp(e.g., subject 1) andtarget(e.g., dog), calculate the mean and SD ofrt - Check if each individual
rtfalls withinmean ± 1.96*SD - Replace any out-of-range values with
NA
Looking at your sample data, you'll see that subject 3's W group dog value (4002) gets replaced with NA—which makes sense, since it's way outside the range for that specific subject-target group.
内容的提问来源于stack exchange,提问作者Jen

