如何使用R将数据集划分为两个训练集和一个测试集?
Hey there! I get that you want to split your dataset into two training sets and one test set in R—let me break this down for you with simple, reproducible steps. The core idea is to first split off your test set, then split the remaining data into your two training subsets.
方法一:基础R实现(无需额外包)
If you don't want to install any extra packages, base R's sample() function works perfectly. Here's a step-by-step example:
# 设置随机种子,确保结果可重复(随便选个数字就行) set.seed(123) # 假设你的数据集叫df # 第一步:拆分出测试集(比如取30%的数据作为测试集) test_indices <- sample(nrow(df), size = 0.3 * nrow(df)) test_set <- df[test_indices, ] train_total <- df[-test_indices, ] # 剩下的70%作为总训练集 # 第二步:把总训练集拆成两个训练集(比如各占总训练集的50%) train1_indices <- sample(nrow(train_total), size = 0.5 * nrow(train_total)) train_set1 <- train_total[train1_indices, ] train_set2 <- train_total[-train1_indices, ]
You can tweak the percentages to fit your needs—for example, if you want a 60%/20%/20% split (train1/train2/test), just adjust the size values accordingly.
方法二:用caret包(推荐分层抽样)
If you're working on a classification task (or want to preserve the distribution of your target variable), the caret package's createDataPartition() function is way better—it does stratified sampling to make sure each subset has the same proportion of target classes as the original data.
# 先安装并加载caret(如果还没装的话) # install.packages("caret") library(caret) set.seed(123) # 假设你的目标变量是y(分类或回归都适用) # 第一步:分层拆分测试集(30%数据) test_partition <- createDataPartition(df$y, p = 0.3, list = FALSE) test_set <- df[test_partition, ] train_total <- df[-test_partition, ] # 第二步:分层拆分总训练集为两个训练集(各50%) train1_partition <- createDataPartition(train_total$y, p = 0.5, list = FALSE) train_set1 <- train_total[train1_partition, ] train_set2 <- train_total[-train1_partition, ]
关键注意事项
- Always use
set.seed(): This ensures your split is reproducible—anyone running your code will get the exact same subsets, which is crucial for debugging and sharing work. - Adjust proportions as needed: No need to stick to 70%/15%/15% or 60%/20%/20%—tweak the
size(base R) orp(caret) values to match your project's requirements. - For regression tasks:
createDataPartition()still works great—it splits based on quantiles of the target variable to preserve its distribution across subsets.
内容的提问来源于stack exchange,提问作者Alison

