You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R包gtsummary构建单变量与多变量coxph模型时add_global_p函数报错的技术咨询

解决gtsummary中add_global_p()报错:Model has aliased term(s); df ambiguous

Hey there, let's break down this error and fix it step by step. The core issue here is perfect multicollinearity (aliased terms) in your survival models—meaning some variables in your model are completely dependent on others, so the model can't distinguish their individual effects. This is what's causing car::Anova() (which add_global_p() relies on under the hood) to throw that "df ambiguous" error.

1. The Most Likely Culprit: Pathological Stage & T/N/M Variables

Looking at your variable list, you're including both Path_stage (pathological stage) and Path_T/Path_N/Path_M (individual T/N/M categories) in your models. Pathological stage is directly derived from combining T, N, and M classifications—they're not independent. Putting all of them in the same model creates a perfect linear dependency, which statistical functions can't handle.

Fix for Multivariable Model

Your multivariable model uses ~. to include all variables, which pulls in both the stage and individual T/N/M variables. You need to pick one set or the other:

Option 1: Keep T/N/M, remove Path_stage

library(survival)
library(gtsummary)

# Explicitly list variables without Path_stage
tbl_mvsurv <- coxph(Surv(Time, Status) ~ Age + Gender + Path_T + Path_N + Path_M + TP53, data = data1) %>% 
  tbl_regression(exponentiate = TRUE) %>% 
  add_global_p() %>% 
  bold_p(t = 0.05) %>% 
  bold_labels()

Option 2: Keep Path_stage, remove T/N/M

# Use Path_stage instead of individual T/N/M variables
tbl_mvsurv <- coxph(Surv(Time, Status) ~ Age + Gender + Path_stage + TP53, data = data1) %>% 
  tbl_regression(exponentiate = TRUE) %>% 
  add_global_p() %>% 
  bold_p(t = 0.05) %>% 
  bold_labels()

2. Tweaks for Your Univariable Model

Your univariable model uses select(data1, everything()), which includes all variables (including both stage and T/N/M). While univariable models test each variable one at a time (so no cross-variable multicollinearity here), you might still run into issues if a variable has a single level (e.g., a categorical variable where all patients fall into one group) or missing levels after na.omit().

Optimize the Univariable Code

Explicitly select only the variables you need to analyze, and double-check each variable's levels:

library(survival)
library(gtsummary)

# Select only relevant variables to avoid clutter
tbl_uvsurv <- data1 %>%
  select(Age, Gender, Path_T, Path_N, Path_M, Path_stage, TP53, Time, Status) %>%
  tbl_uvregression(
    method = coxph,
    y = Surv(Time, Status),
    exponentiate = TRUE
  ) %>%
  add_global_p(include = c("Age", "Gender", "Path_T", "Path_N", "Path_M", "Path_stage", "TP53")) %>%
  bold_p(t = 0.05) %>%
  bold_labels()

# Check for variables with single levels (which can cause issues)
lapply(c("Path_T", "Path_N", "Path_M", "Path_stage"), function(var) table(data1[[var]]))

3. How to Diagnose Hidden Aliased Terms

If you're still seeing errors after fixing the stage/T/N/M issue, use the alias() function to identify exactly which variables are causing collinearity:

# Fit a temporary model to check for aliased terms
temp_model <- coxph(Surv(Time, Status) ~ ., data = data1)
alias(temp_model)

This will show you which variables are redundant and can be removed.


内容的提问来源于stack exchange,提问作者belivemyself96

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 11:52:31