You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按患者分组拆分数据集并维持LesionResponse类别均衡的方案咨询

保持患者完整性的数据集分层抽样解决方案(针对LesionResponse预测模型)

问题描述

我要构建LesionResponse预测模型,计划按60%训练集、20%验证集、20%测试集拆分数据集。但数据集中同一PatientID对应多行数据,必须保证同一患者的所有数据归到同一集合以避免偏差。目前已实现按患者分组拆分,但各集合的LesionResponse类别分布和原数据集的70%(1)、30%(0)不符,训练集当前为55/45,需要修正该问题。

当前集合类别分布统计

summary(train$LesionResponse)
#   0   1
# 159 487
summary(validation$LesionResponse)
#  0   1
# 33 170
summary(test$LesionResponse)
#  0   1
# 77 126

数据集结构

structure(list(PatientID = c("P1", "P1", "P1", 
"P2", "P3", "P3", "P4", "P5", 
"P5", "P6"), LesionResponse = structure(c(2L, 2L, 1L, 2L, 2L, 2L, 2L, 
    2L, 1L, 2L), .Label = c("0", 
    "1"), class = "factor"), pyrad_tum_original_shape_LeastAxisLength = c(19.7842995242803, 
    15.0703960571122, 21.0652247652897, 11.804125918871, 27.3980336338908, 
    17.0584330264122, 4.90406343942677, 4.78480430022189, 6.2170232078547, 
    5.96309532740722, 5.30141540007441), pyrad_tum_original_shape_Sphericity = c(0.652056853392657, 
    0.773719977240238, 0.723869070051882, 0.715122964970338, 
    0.70796498824535, 0.811937882810929, 0.836458991713367, 0.863337931630415, 
    0.851654860256904, 0.746212862162174), pyrad_tum_log.sigma.5.0.mm.3D_firstorder_Skewness = c(0.367453961973625, 
    0.117673346718817, 0.0992025164349288, -0.174029385779302, 
    -0.863570016875989, -0.8482193060411, -0.425424618080682, 
    -0.492420174157913, 0.0105111292451967, 0.249865833210199), pyrad_tum_log.sigma.5.0.mm.3D_glcm_Contrast = c(0.376932105256115, 
    0.54885738172596, 0.267158344601612, 2.90094719958076, 0.322424096161189, 
    0.221356030145403, 1.90012334870722, 0.971638740404501, 0.31547550396399, 
    0.653999340294952), pyrad_tum_wavelet.LHH_glszm_GrayLevelNonUniformityNormalized = c(0.154973213866752, 
    0.176128379241556, 0.171129002059539, 0.218343919352019, 
    0.345985943932352, 0.164905080489496, 0.104536489151874, 
    0.1280276816609, 0.137912385073012, 0.133420904484894), pyrad_tum_wavelet.LHH_glszm_LargeAreaEmphasis = c(27390.2818110851, 
    11327.7931034483, 51566.7948885976, 7261.68702290076, 340383.536555142, 
    22724.7792207792, 45.974358974359, 142.588235294118, 266.744186046512, 
    1073.45205479452), pyrad_tum_wavelet.LHH_glszm_LargeAreaLowGrayLevelEmphasis = c(677.011907073653, 
    275.281153810458, 582.131636238695, 173.747506476692, 6140.73990175018, 
    558.277670638306, 1.81042257642817, 4.55724031114589, 6.51794350173746, 
    19.144924585586), pyrad_tum_wavelet.LHH_glszm_SizeZoneNonUniformityNormalized = c(0.411899490603372, 
    0.339216399209913, 0.425584323452468, 0.355165782879786, 
    0.294934042125209, 0.339208410636982, 0.351742274819198, 
    0.394463667820069, 0.360735532720389, 0.36911240382811)), row.names = c(NA, -10L), class = c("tbl_df", "tbl", 
"data.frame"))

我曾考虑循环拆分unique(PatientID)数据集,若类别分布失衡则重复拆分,但希望找到更优方案。


解决方案:患者级分层抽样

核心思路是先按患者聚合确定标签,再对患者集合做分层抽样,确保类别分布匹配原数据集,同时保证同一患者数据归属同一集合。

步骤1:构建患者-标签映射表

先为每个PatientID确定对应的LesionResponse类别(若一个患者同时有0和1记录,可按多数类别/临床规则定义,以下以多数类别为例):

library(dplyr)

# 生成患者-标签映射
patient_labels <- df %>%
  group_by(PatientID) %>%
  summarise(LesionResponse = names(which.max(table(LesionResponse))),
            .groups = "drop")

步骤2:分层拆分患者集合

按60/20/20比例,以LesionResponse为分层变量拆分患者:

set.seed(123) # 固定随机种子保证结果可复现

# 抽取训练集患者(60%)
train_patients <- patient_labels %>%
  group_by(LesionResponse) %>%
  sample_frac(0.6) %>%
  ungroup()

# 剩余患者集合
remaining_patients <- anti_join(patient_labels, train_patients, by = "PatientID")

# 从剩余中抽取验证集患者(50%,即总数据的20%)
val_patients <- remaining_patients %>%
  group_by(LesionResponse) %>%
  sample_frac(0.5) %>%
  ungroup()

# 剩余为测试集患者
test_patients <- anti_join(remaining_patients, val_patients, by = "PatientID")

步骤3:映射回原始数据集

根据患者分组拆分原始数据:

train <- df %>% filter(PatientID %in% train_patients$PatientID)
validation <- df %>% filter(PatientID %in% val_patients$PatientID)
test <- df %>% filter(PatientID %in% test_patients$PatientID)

进阶优化方案

  • 加权分层抽样:若不同患者的样本行数差异大,可给患者赋予样本量权重,做加权分层抽样,同时保证集合的样本量比例和类别分布匹配。
  • 使用caret包简化操作:利用caret包的createDataPartition函数快速生成分层索引:
library(caret)

# 生成训练集患者索引
patient_indices <- createDataPartition(patient_labels$LesionResponse, p = 0.6, list = FALSE)
train_patients <- patient_labels[patient_indices, ]

# 后续拆分验证集、测试集步骤同前

内容的提问来源于stack exchange,提问作者NDe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 05:48:24