You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过循环从同一数据集生成不同大小的独立样本数据集?

问题分析与解决方案

原代码有两个核心问题导致无法得到预期的三个数据集:

  1. 循环中每次都重新赋值results变量,之前的结果会被直接覆盖;
  2. results[,i] = sample_data是错误的赋值方式——i是样本的行数,不是列索引,这会导致列位置越界或赋值失败。

以下是两种可行的解决方式:

方式一:创建独立命名的DataFrame

直接用assign()函数根据样本量生成对应变量名,把采样结果赋值给这些变量:

n.values <- c(500, 1000, 2000)
# 加载dplyr包,或用dplyr::sample_n明确指定函数
library(dplyr)
for (i in n.values) {
  sample_data <- sample_n(train, i)
  # 生成目标变量名,比如results500
  var_name <- paste0("results", i)
  # 将采样数据赋值给对应变量
  assign(var_name, sample_data)
}

运行后环境中会生成results500、results1000、results2000三个独立的DataFrame,每个对应指定大小的样本。

方式二:用列表统一管理(更推荐)

在R中,用列表管理多个同类型对象会更规范,避免环境中变量泛滥,也方便后续批量操作:

n.values <- c(500, 1000, 2000)
library(dplyr)
results_list <- list()
for (i in n.values) {
  sample_data <- sample_n(train, i)
  # 把采样数据存入列表,用对应的名称做索引
  results_list[[paste0("results", i)]] <- sample_data
}
# 调用单个数据集时用:results_list[["results500"]]

内容的提问来源于stack exchange,提问作者Lola Ro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 14:45:57