如何在R中模拟随时间变化且组内相关的RT响应数据
问题:如何模拟符合时间变化规律的反应时间(RT)数据
我正在使用R进行功效分析的数据模拟,旨在研究实验场景中儿童的反应时间(RT)对其后续语言能力的预测作用,儿童将按分组变量Group划分。拟采用的统计模型为:language ~ RT * Timepoint + Group。我需要模拟RT变量,使其随Timepoint呈预期方向变化:
- 总体随时间增加约0.2
- 个体间变化差异的标准差约0.1
- 同一受试者内的RT具有相关性
目前已有如下代码,但生成的RT变量未呈现受试者随时间点的预期变化规律,请问如何修改代码实现需求?
# First specify the values for each variable in the model SimulatedData <- function( MeanIntrcpt = 37.5, # baseline language score MeanSlope = 14.63, # overall effect of group on language score Group = 0.39, # estimated effect of group on language score time_slope = .2, # estimated effect of Timepoint on reaction time n_subjects = 60, # number of subjects n_timepoints = 3, # number of timepoints RandIntrcpt_subjects = 0.2, # Random intercept for subjects RandSlope_subjects = 0.2, # Random slope for subjects Randslope_subject_time = 0.1 # Jittered effect of time between subjects ){ # Simulate data to reflect our subjects SimulatedSubjects <- faux::rnorm_multi( n = n_subjects, mu = 0, sd = c(RandIntrcpt_subjects, RandSlope_subjects, Randslope_subject_time), varnames = c("RandIntrcpt_subjects_StdDev", "RandSlope_subjects_StdDev", "RandSlope_subjects_time_StdDev")) %>% mutate(ChildID = faux::make_id(nrow(.), "S_"), Group = sample(c(1, 2), 60, replace= T)) # assign half the participants to group 1/group 2 # Simulate data to reflect timepoints timepoints <- data.frame(Timepoint = seq_len(n_timepoints)) # create the combined data frame - simulated subjects crossed with timepoints SimulatedSubjectsAndStimuli <- crossing(SimulatedSubjects, timepoints) # Simulate RT data, 60 subjects * 3 timepoints = 180 rows SimulatedSubjectsAndStimuli <- SimulatedSubjectsAndStimuli %>% mutate(RT = runif(n_subjects*n_timepoints, min = -0.5, max = 0.5)) } SimulatedData <- SimulatedData() SimulatedData <- SimulatedData() > head(SimulatedData) # A tibble: 6 × 7 RandIntrcpt_subjects_StdDev RandSlope_subjects_StdDev RandSlope_subjects_time_StdDev ChildID Group Timepoint RT <dbl> <dbl> <dbl> <chr> <dbl> <int> <dbl> 1 -0.584 0.0153 -0.137 S_42 1 1 -0.452 2 -0.584 0.0153 -0.137 S_42 1 2 -0.265 3 -0.584 0.0153 -0.137 S_42 1 3 -0.0435 4 -0.412 -0.157 0.169 S_15 1 1 -0.491 5 -0.412 -0.157 0.169 S_15 1 2 -0.176 6 -0.412 -0.157 0.169 S_15 1 3 0.152 >
修改方案
原代码的核心问题是RT仅通过runif生成完全随机的数值,未结合时间趋势和个体随机效应,因此无法呈现预期的时间变化规律和个体内相关性。以下是修改后的完整代码:
# 定义数据模拟函数 SimulatedData <- function( MeanIntrcpt = 37.5, # 语言能力基线得分 MeanSlope = 14.63, # 分组对语言能力的总体效应 GroupEffect = 0.39, # 分组对语言能力的估计效应 time_slope = 0.2, # RT随时间变化的总体斜率(每时间点增加0.2) n_subjects = 60, # 受试者数量 n_timepoints = 3, # 时间点数量 RandIntrcpt_subjects = 0.2, # 受试者RT的随机截距标准差 RandSlope_subject_time = 0.1,# 受试者RT随时间变化的随机斜率标准差 RT_residual_sd = 0.05 # RT的残差标准差(可根据需求调整) ){ # 生成受试者水平的随机效应 SimulatedSubjects <- faux::rnorm_multi( n = n_subjects, mu = 0, sd = c(RandIntrcpt_subjects, RandSlope_subject_time), varnames = c("RT_RandInt", "RT_RandSlope") ) %>% mutate( ChildID = faux::make_id(nrow(.), "S_"), Group = sample(c(1, 2), n_subjects, replace = TRUE) # 随机分配分组 ) # 生成时间点数据 timepoints <- data.frame(Timepoint = seq_len(n_timepoints)) # 交叉受试者和时间点,构建长格式数据 SimulatedData_long <- tidyr::crossing(SimulatedSubjects, timepoints) # 模拟RT数据:结合固定效应、个体随机效应和残差 SimulatedData_long <- SimulatedData_long %>% mutate( RT = # 固定效应:总体随时间增加0.2 time_slope * Timepoint + # 个体随机截距:每个受试者的基线RT差异 RT_RandInt + # 个体随机斜率:每个受试者的时间变化差异(标准差0.1) RT_RandSlope * Timepoint + # 残差项:添加少量随机噪声 rnorm(nrow(.), mean = 0, sd = RT_residual_sd) ) # 返回模拟数据 return(SimulatedData_long) } # 生成模拟数据 SimulatedData <- SimulatedData() # 查看前几行数据 head(SimulatedData)
关键修改说明
RT生成逻辑重构:
- 加入
time_slope * Timepoint实现总体时间趋势:每增加一个时间点,RT平均上升0.2 - 加入
RT_RandInt控制个体基线RT差异:每个受试者有独特的初始RT水平 - 加入
RT_RandSlope * Timepoint实现个体间时间变化差异:不同受试者的RT随时间变化的速率不同,标准差由RandSlope_subject_time = 0.1控制 - 加入残差项
rnorm(...):在个体趋势基础上添加少量随机噪声,更贴近真实数据
- 加入
函数优化:
- 修正原函数未返回数据的问题,添加
return(SimulatedData_long) - 简化随机效应变量名,提升代码可读性
- 分组抽样时使用
n_subjects替代硬编码的60,保证参数一致性
- 修正原函数未返回数据的问题,添加
个体内相关性实现:
同一受试者的RT共享相同的随机截距和随机斜率,因此同一受试者在不同时间点的RT会呈现自然的相关性,无需额外指定协方差结构。
内容的提问来源于stack exchange,提问作者Catherine Laing
相关产品推荐
相关产品推荐

