如何循环运行pmsampsize函数批量计算样本量与结局事件数?
背景示例
运行以下R代码可得到单个参数下的样本量计算结果:
library(pmsampsize) pmsampsize(type = "s", csrsquared = 0.5, parameters = 10, rate = 0.065, timepoint = 2, meanfup = 2.07)
输出结果:
NB: Assuming 0.05 acceptable difference in apparent & adjusted R-squared NB: Assuming 0.05 margin of error in estimation of overall risk at time point = 2 NB: Events per Predictor Parameter (EPP) assumes overall event rate = 0.065 Samp_size Shrinkage Parameter CS_Rsq Max_Rsq Nag_Rsq EPP Criteria 1 5143 0.900 30 0.051 0.555 0.092 23.07 Criteria 2 1039 0.648 30 0.051 0.555 0.092 4.66 Criteria 3 * 5143 0.900 30 0.051 0.555 0.092 23.07 Final SS 5143 0.900 30 0.051 0.555 0.092 23.07 Minimum sample size required for new model development based on user inputs = 5143, corresponding to 10646 person-time** of follow-up, with 692 outcome events assuming an overall event rate = 0.065 and therefore an EPP = 23.07
需求说明
需要通过循环函数,针对csrsquared(取值序列R=seq(from=0, to=0.5, by=0.1))和parameters(取值序列PA=seq(from=1, to=10, by=1))的所有组合,批量运行pmsampsize函数,计算对应的最小样本量和结局事件数,并整理成如下格式的表格:
| csrsquared (R) | parameters (PA) | Minimum sample size | outcome events |
|---|
示例结构:
R PA minimum sample size outcome events 0 1 0 2 0 3 0 4
问题代码
编写了如下代码但无法正常运行,请求帮助修改:
library(tidyverse) library (pmsampsize) foo = function(R,PA){ T=pmsampsize(type = "s", csrsquared = R, parameters = PA, rate = 0.065,timepoint = 2, meanfup = 2.07) T$events T$sample_size } df = tibble( R=seq(from = 0, to = 0.5,by=0.1) PA=seq(from = 0, to = 10,by=1) )
修正后的代码及解释
原代码存在3个核心问题:自定义函数仅返回单个值、参数组合未生成全量、数据框创建语法错误。修正后的代码如下:
library(tidyverse) library(pmsampsize) # 定义函数,同时返回样本量和结局事件数 calc_sample_events <- function(R, PA) { result <- pmsampsize(type = "s", csrsquared = R, parameters = PA, rate = 0.065, timepoint = 2, meanfup = 2.07) tibble( minimum_sample_size = result$sample_size, outcome_events = result$events ) } # 生成R和PA的所有参数组合 df <- expand_grid( R = seq(from = 0, to = 0.5, by = 0.1), PA = seq(from = 1, to = 10, by = 1) ) # 批量计算并合并结果 final_df <- df %>% rowwise() %>% mutate(calc_sample_events(R, PA)) %>% ungroup() # 查看最终表格 print(final_df)
关键修正点:
- 用
expand_grid生成R和PA的全量组合,解决原代码未覆盖所有参数配对的问题; - 自定义函数返回包含两个目标指标的
tibble,确保能同时获取样本量和事件数; - 通过
rowwise()+mutate()实现逐行调用计算函数,自动将结果合并到原数据框; - 修正原数据框创建时的语法错误,同时将
PA的起始值调整为1,匹配需求中的参数序列。
运行代码后,final_df即为所需的完整结果表格。
内容的提问来源于stack exchange,提问作者elisa
相关产品推荐
相关产品推荐

