R语言survminer包ggsurvplot函数newdata参数使用报错问题
解决ggsurvplot绘制生存曲线的两类legend.labs报错问题
问题1:使用newdata时提示legend.labs长度应为观测数
原因
survfit(cox_model, newdata = d.dat.tot.clean)会为数据框的每一行观测生成一条生存曲线,而非按基因型分组生成曲线。ggsurvplot识别到曲线数量等于观测数(你遇到的是236),但传入的legend.labs是因子水平数(3个),长度不匹配导致报错。
解决办法
newdata需要传入仅包含各因子水平的数据集,而非原始全量数据。用expand.grid生成每个基因型的代表行,确保survfit只生成对应分组的3条曲线:
set.seed(123) d.dat.tot <- data.frame( DONOR_SURVIVAL_TIME = rexp(100, 0.1), DEAD_OR_ALIVE = sample(0:1, 100, replace = TRUE), GENOTYPE.111758446 = sample(c("A/A", "G/A", "G/G"), 100, replace = TRUE) ) d.dat.tot.clean <- d.dat.tot[d.dat.tot$DEAD_OR_ALIVE < 2, ] d.dat.tot.clean$GENOTYPE.111758446 <- as.factor(d.dat.tot.clean$GENOTYPE.111758446) cox_model_single_111758446 <- coxph( Surv(DONOR_SURVIVAL_TIME, DEAD_OR_ALIVE) ~ GENOTYPE.111758446, data = d.dat.tot.clean ) # 生成仅包含各基因型水平的newdata newdata_grouped <- expand.grid(GENOTYPE.111758446 = levels(d.dat.tot.clean$GENOTYPE.111758446)) cox_plot_single_111758446 <- ggsurvplot( survfit(cox_model_single_111758446, newdata = newdata_grouped), data = d.dat.tot.clean, pval = TRUE, conf.int = TRUE, risk.table = TRUE, legend.title = "Genotype", legend.labs = levels(d.dat.tot.clean$GENOTYPE.111758446), xlab = "Time to Last Follow-up, mo", ylab = "Cumulative Survival, %", ggtheme = theme_minimal(), palette = "set2" ) print(cox_plot_single_111758446)
问题2:不使用newdata时提示legend.labs长度应为1
原因
直接调用survfit(cox_model_single_111758446)默认生成的是基准风险的单条生存曲线,而非按基因型分组的多条曲线。ggsurvplot识别到只有1条曲线,但传入了3个legend.labs,长度不匹配报错。
解决办法
方案1:直接按分组拟合survfit(无需依赖cox模型)
set.seed(123) d.dat.tot <- data.frame( DONOR_SURVIVAL_TIME = rexp(100, 0.1), DEAD_OR_ALIVE = sample(0:1, 100, replace = TRUE), GENOTYPE.111758446 = sample(c("A/A", "G/A", "G/G"), 100, replace = TRUE) ) d.dat.tot.clean <- d.dat.tot[d.dat.tot$DEAD_OR_ALIVE < 2, ] d.dat.tot.clean$GENOTYPE.111758446 <- as.factor(d.dat.tot.clean$GENOTYPE.111758446) # 直接按基因型分组拟合survfit surv_fit <- survfit(Surv(DONOR_SURVIVAL_TIME, DEAD_OR_ALIVE) ~ GENOTYPE.111758446, data = d.dat.tot.clean) cox_plot_single_111758446 <- ggsurvplot( surv_fit, data = d.dat.tot.clean, pval = TRUE, conf.int = TRUE, risk.table = TRUE, legend.title = "Genotype", legend.labs = levels(d.dat.tot.clean$GENOTYPE.111758446), xlab = "Time to Last Follow-up, mo", ylab = "Cumulative Survival, %", ggtheme = theme_minimal(), palette = "set2" ) print(cox_plot_single_111758446)
方案2:依然基于cox模型,传入分组后的newdata(同问题1的解决逻辑)
内容的提问来源于stack exchange,提问作者Yukiyaama
相关产品推荐
相关产品推荐

