如何在R语言中非随机拆分训练集与测试集(取前80%为训练集)
Fixing Non-Random Train/Test Split in R
Hey there! Let's work through fixing your code to split your data into a training set (first 80% of rows) and test set (remaining 20%) correctly.
First, let's spot the small issues in your original code:
- Typos: You defined
rows = nrow(data)but then used0.80*row(missing an 's' at the end ofrow) - Non-integer index: Calculating
0.8 * rowsmight give a decimal number, but R requires integer values for row indexing—we need to convert that to a whole number - Incorrect subset syntax:
data[1:index]selects columns instead of rows. To grab rows, you need to include a comma to indicate "all columns":data[1:index, ]
Corrected Base R Code
Here's the fixed version that addresses all these issues:
# Get total number of rows in your dataset total_rows <- nrow(data) # Calculate the last row index for the training set (use floor() to get an integer) train_last_row <- floor(0.8 * total_rows) # Split the data train_data <- data[1:train_last_row, ] # First 80% of rows, all columns test_data <- data[(train_last_row + 1):total_rows, ] # Remaining rows, all columns
Alternative with Tidyverse (dplyr)
If you prefer using the tidyverse ecosystem, the slice() function makes this even more readable:
library(dplyr) # Calculate the split index (same as before) train_last_row <- floor(0.8 * nrow(data)) # Split using slice() train_data <- data %>% slice(1:train_last_row) test_data <- data %>% slice((train_last_row + 1):n()) # n() gets the total number of rows
Just remember to replace data with the actual name of your data frame!
内容的提问来源于stack exchange,提问作者user15051990
相关产品推荐
相关产品推荐

