You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用R将数据集划分为两个训练集和一个测试集?

拆分数据集为两个训练集+一个测试集的R实现

Hey there! I get that you want to split your dataset into two training sets and one test set in R—let me break this down for you with simple, reproducible steps. The core idea is to first split off your test set, then split the remaining data into your two training subsets.

方法一:基础R实现(无需额外包)

If you don't want to install any extra packages, base R's sample() function works perfectly. Here's a step-by-step example:

# 设置随机种子,确保结果可重复(随便选个数字就行)
set.seed(123)

# 假设你的数据集叫df
# 第一步:拆分出测试集(比如取30%的数据作为测试集)
test_indices <- sample(nrow(df), size = 0.3 * nrow(df))
test_set <- df[test_indices, ]
train_total <- df[-test_indices, ]  # 剩下的70%作为总训练集

# 第二步:把总训练集拆成两个训练集(比如各占总训练集的50%)
train1_indices <- sample(nrow(train_total), size = 0.5 * nrow(train_total))
train_set1 <- train_total[train1_indices, ]
train_set2 <- train_total[-train1_indices, ]

You can tweak the percentages to fit your needs—for example, if you want a 60%/20%/20% split (train1/train2/test), just adjust the size values accordingly.

方法二:用caret包(推荐分层抽样)

If you're working on a classification task (or want to preserve the distribution of your target variable), the caret package's createDataPartition() function is way better—it does stratified sampling to make sure each subset has the same proportion of target classes as the original data.

# 先安装并加载caret(如果还没装的话)
# install.packages("caret")
library(caret)

set.seed(123)

# 假设你的目标变量是y(分类或回归都适用)
# 第一步:分层拆分测试集(30%数据)
test_partition <- createDataPartition(df$y, p = 0.3, list = FALSE)
test_set <- df[test_partition, ]
train_total <- df[-test_partition, ]

# 第二步:分层拆分总训练集为两个训练集(各50%)
train1_partition <- createDataPartition(train_total$y, p = 0.5, list = FALSE)
train_set1 <- train_total[train1_partition, ]
train_set2 <- train_total[-train1_partition, ]

关键注意事项

  • Always use set.seed(): This ensures your split is reproducible—anyone running your code will get the exact same subsets, which is crucial for debugging and sharing work.
  • Adjust proportions as needed: No need to stick to 70%/15%/15% or 60%/20%/20%—tweak the size (base R) or p (caret) values to match your project's requirements.
  • For regression tasks: createDataPartition() still works great—it splits based on quantiles of the target variable to preserve its distribution across subsets.

内容的提问来源于stack exchange,提问作者Alison

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:26:04