You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中按位置拆分数据集:奇数行入训练集、整十行入测试集

Split Your Dataset into Training & Test Sets (R)

Got it, let's walk through how to split your 17,000-row dataset exactly as you need it—since you're working with a data.frame in R, this is straightforward to pull off.

First, let's assume your original dataset is stored in a data.frame named original_data. Here's the code to split it based on your rules:

# Grab the total number of rows in your dataset
total_rows <- nrow(original_data)

# Create training set: all odd-positioned instances (1, 3, 5, ...)
train_data <- original_data[seq(1, total_rows, by = 2), ]

# Create test set: instances at positions 10, 20, 30, ...
test_data <- original_data[seq(10, total_rows, by = 10), ]

A Quick Breakdown:

  • seq(1, total_rows, by = 2) generates a sequence starting at 1, ending at the last row number, and jumping by 2 each time. This directly gives you all the odd indices you need for the training set.
  • seq(10, total_rows, by = 10) starts at 10 and increments by 10, so you get every 10th row starting from the 10th position—perfect for your test set.

Verify the Results (Optional):

If you want to double-check that the splits are correct, run these lines to confirm row counts:

# Check how many rows are in the training set
cat("Training set row count:", nrow(train_data), "\n")

# Check how many rows are in the test set
cat("Test set row count:", nrow(test_data), "\n")

For your 17,000-row dataset:

  • The training set should have 8,500 rows (half of 17k, since we're taking every odd row)
  • The test set should have 1,700 rows (17k divided by 10, since we're taking every 10th row starting at 10)

This method is efficient, easy to read, and aligns perfectly with your exact requirements. No extra hoops to jump through—just clean, direct code.

内容的提问来源于stack exchange,提问作者Rahul Rai Khatri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:25:42