按患者分组拆分数据集并维持LesionResponse类别均衡的方案咨询
保持患者完整性的数据集分层抽样解决方案(针对LesionResponse预测模型)
问题描述
我要构建LesionResponse预测模型,计划按60%训练集、20%验证集、20%测试集拆分数据集。但数据集中同一PatientID对应多行数据,必须保证同一患者的所有数据归到同一集合以避免偏差。目前已实现按患者分组拆分,但各集合的LesionResponse类别分布和原数据集的70%(1)、30%(0)不符,训练集当前为55/45,需要修正该问题。
当前集合类别分布统计
summary(train$LesionResponse) # 0 1 # 159 487 summary(validation$LesionResponse) # 0 1 # 33 170 summary(test$LesionResponse) # 0 1 # 77 126
数据集结构
structure(list(PatientID = c("P1", "P1", "P1", "P2", "P3", "P3", "P4", "P5", "P5", "P6"), LesionResponse = structure(c(2L, 2L, 1L, 2L, 2L, 2L, 2L, 2L, 1L, 2L), .Label = c("0", "1"), class = "factor"), pyrad_tum_original_shape_LeastAxisLength = c(19.7842995242803, 15.0703960571122, 21.0652247652897, 11.804125918871, 27.3980336338908, 17.0584330264122, 4.90406343942677, 4.78480430022189, 6.2170232078547, 5.96309532740722, 5.30141540007441), pyrad_tum_original_shape_Sphericity = c(0.652056853392657, 0.773719977240238, 0.723869070051882, 0.715122964970338, 0.70796498824535, 0.811937882810929, 0.836458991713367, 0.863337931630415, 0.851654860256904, 0.746212862162174), pyrad_tum_log.sigma.5.0.mm.3D_firstorder_Skewness = c(0.367453961973625, 0.117673346718817, 0.0992025164349288, -0.174029385779302, -0.863570016875989, -0.8482193060411, -0.425424618080682, -0.492420174157913, 0.0105111292451967, 0.249865833210199), pyrad_tum_log.sigma.5.0.mm.3D_glcm_Contrast = c(0.376932105256115, 0.54885738172596, 0.267158344601612, 2.90094719958076, 0.322424096161189, 0.221356030145403, 1.90012334870722, 0.971638740404501, 0.31547550396399, 0.653999340294952), pyrad_tum_wavelet.LHH_glszm_GrayLevelNonUniformityNormalized = c(0.154973213866752, 0.176128379241556, 0.171129002059539, 0.218343919352019, 0.345985943932352, 0.164905080489496, 0.104536489151874, 0.1280276816609, 0.137912385073012, 0.133420904484894), pyrad_tum_wavelet.LHH_glszm_LargeAreaEmphasis = c(27390.2818110851, 11327.7931034483, 51566.7948885976, 7261.68702290076, 340383.536555142, 22724.7792207792, 45.974358974359, 142.588235294118, 266.744186046512, 1073.45205479452), pyrad_tum_wavelet.LHH_glszm_LargeAreaLowGrayLevelEmphasis = c(677.011907073653, 275.281153810458, 582.131636238695, 173.747506476692, 6140.73990175018, 558.277670638306, 1.81042257642817, 4.55724031114589, 6.51794350173746, 19.144924585586), pyrad_tum_wavelet.LHH_glszm_SizeZoneNonUniformityNormalized = c(0.411899490603372, 0.339216399209913, 0.425584323452468, 0.355165782879786, 0.294934042125209, 0.339208410636982, 0.351742274819198, 0.394463667820069, 0.360735532720389, 0.36911240382811)), row.names = c(NA, -10L), class = c("tbl_df", "tbl", "data.frame"))
我曾考虑循环拆分unique(PatientID)数据集,若类别分布失衡则重复拆分,但希望找到更优方案。
解决方案:患者级分层抽样
核心思路是先按患者聚合确定标签,再对患者集合做分层抽样,确保类别分布匹配原数据集,同时保证同一患者数据归属同一集合。
步骤1:构建患者-标签映射表
先为每个PatientID确定对应的LesionResponse类别(若一个患者同时有0和1记录,可按多数类别/临床规则定义,以下以多数类别为例):
library(dplyr) # 生成患者-标签映射 patient_labels <- df %>% group_by(PatientID) %>% summarise(LesionResponse = names(which.max(table(LesionResponse))), .groups = "drop")
步骤2:分层拆分患者集合
按60/20/20比例,以LesionResponse为分层变量拆分患者:
set.seed(123) # 固定随机种子保证结果可复现 # 抽取训练集患者(60%) train_patients <- patient_labels %>% group_by(LesionResponse) %>% sample_frac(0.6) %>% ungroup() # 剩余患者集合 remaining_patients <- anti_join(patient_labels, train_patients, by = "PatientID") # 从剩余中抽取验证集患者(50%,即总数据的20%) val_patients <- remaining_patients %>% group_by(LesionResponse) %>% sample_frac(0.5) %>% ungroup() # 剩余为测试集患者 test_patients <- anti_join(remaining_patients, val_patients, by = "PatientID")
步骤3:映射回原始数据集
根据患者分组拆分原始数据:
train <- df %>% filter(PatientID %in% train_patients$PatientID) validation <- df %>% filter(PatientID %in% val_patients$PatientID) test <- df %>% filter(PatientID %in% test_patients$PatientID)
进阶优化方案
- 加权分层抽样:若不同患者的样本行数差异大,可给患者赋予样本量权重,做加权分层抽样,同时保证集合的样本量比例和类别分布匹配。
- 使用caret包简化操作:利用
caret包的createDataPartition函数快速生成分层索引:
library(caret) # 生成训练集患者索引 patient_indices <- createDataPartition(patient_labels$LesionResponse, p = 0.6, list = FALSE) train_patients <- patient_labels[patient_indices, ] # 后续拆分验证集、测试集步骤同前
内容的提问来源于stack exchange,提问作者NDe
相关产品推荐
相关产品推荐

