You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于ctree的Bootstrapping与递归分区分析技术求助

递归分区与Bootstrapping分析解决方案

核心需求:

  • 从递归分区的终端节点提取3年生存率,将患者按生存率自动分为A(>70%)、B(50%-70%)、C(25%-50%)、D(<25%)组
  • 在Bootstrapping迭代中,记录每位患者每次被分配的组别及最终频次

一、优化单棵决策树的节点分组逻辑

替换原始手动匹配节点的代码,实现自动化节点-生存率映射与分组,无需手动指定节点编号:

library(partykit)
library(survival)
library(dplyr)

# 加载公开数据
data("GBSG2", package = "TH.data")
df <- GBSG2

# 构建递归分区树
stree <- ctree(Surv(time, cens) ~ ., 
               data = df, 
               control = ctree_control(minsplit = 50, alpha = 0.1, multiway = TRUE))

# 标记患者所属终端节点
df$node <- predict(stree, type = "node")

# 计算每个节点的3年生存率(1095天)
surv_fit <- survfit(Surv(time, cens) ~ node, data = df)
surv_3y <- summary(surv_fit, times = 365*3)$surv
node_surv_map <- tibble(node = as.integer(names(surv_3y)), surv_3y = surv_3y)

# 自动划分A/B/C/D组
df <- df %>%
  left_join(node_surv_map, by = "node") %>%
  mutate(grp = case_when(
    surv_3y > 0.7 ~ "A",
    surv_3y >= 0.5 & surv_3y <= 0.7 ~ "B",
    surv_3y >= 0.25 & surv_3y < 0.5 ~ "C",
    surv_3y < 0.25 ~ "D",
    TRUE ~ NA_character_
  ))

# 查看分组结果分布
table(df$grp)

二、Bootstrapping迭代与组别频次统计

实现Bootstrapping循环,每次迭代重抽样构建树、分配组别,最终统计每位患者的组别频次:

# 设置迭代次数(测试用10次,最终可改为1000次)
n_iter <- 10

# 初始化结果存储容器
bootstrap_results <- list()

# 执行Bootstrapping循环
for (i in 1:n_iter) {
  # 有放回重抽样生成训练集
  boot_sample <- df %>% sample_frac(size = 1, replace = TRUE)
  
  # 基于抽样集构建递归分区树
  boot_tree <- ctree(Surv(time, cens) ~ ., 
                     data = boot_sample, 
                     control = ctree_control(minsplit = 50, alpha = 0.1, multiway = TRUE))
  
  # 给原始数据集所有患者分配节点
  boot_nodes <- predict(boot_tree, newdata = df, type = "node")
  
  # 计算抽样集中每个节点的3年生存率
  boot_surv_fit <- survfit(Surv(time, cens) ~ boot_nodes, data = boot_sample)
  boot_surv_3y <- summary(boot_surv_fit, times = 365*3)$surv
  boot_node_surv <- tibble(node = as.integer(names(boot_surv_3y)), surv_3y = boot_surv_3y)
  
  # 生成本次迭代的患者分组结果
  boot_grp_result <- tibble(patient_id = 1:nrow(df), # 给患者添加唯一标识
                            node = boot_nodes) %>%
    left_join(boot_node_surv, by = "node") %>%
    mutate(grp = case_when(
      surv_3y > 0.7 ~ "A",
      surv_3y >= 0.5 & surv_3y <= 0.7 ~ "B",
      surv_3y >= 0.25 & surv_3y < 0.5 ~ "C",
      surv_3y < 0.25 ~ "D",
      TRUE ~ NA_character_
    )) %>%
    select(patient_id, grp) %>%
    mutate(iteration = i)
  
  # 存储本次迭代结果
  bootstrap_results[[i]] <- boot_grp_result
}

# 合并所有迭代结果
all_iter_results <- bind_rows(bootstrap_results)

# 统计每位患者的组别频次
patient_grp_freq <- all_iter_results %>%
  group_by(patient_id, grp) %>%
  count(name = "freq") %>%
  pivot_wider(names_from = grp, values_from = freq, values_fill = 0)

# 查看前5位患者的频次结果
head(patient_grp_freq)

关键说明:

  • 每次迭代用有放回抽样数据集构建树,但对原始全量患者进行分组预测,确保所有患者都被覆盖
  • 节点的3年生存率基于抽样数据集计算,符合Bootstrapping的统计逻辑
  • 最终patient_grp_freq表存储了每位患者在所有迭代中被分到A/B/C/D组的次数

内容的提问来源于stack exchange,提问作者Tobias Berg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 05:05:36