如何在R原DataFrame中标记按年随机抽取的样本?
Hey there! Let's work through this problem—since you're handling large biological archival data, we need a clean, memory-efficient way to mark your selected rows directly in the original DataFrame, no more messy duplicates or Excel workarounds.
Core Idea
Instead of extracting subsets and merging back (which causes duplicate issues), we'll add a marker column directly to your original dataset by grouping samples by year, then randomly selecting the required number of entries per group and tagging them.
Solution 1: Using dplyr (Clean & Readable)
This is great for code clarity, and it’s efficient even for large datasets like yours.
First, let’s adapt your sample code to mark selected rows:
# Set seed for reproducible results (critical for research!) set.seed(123) # Your sample dataset dat <- data.frame( X = sample(2000:2016, 50, replace=TRUE), Y = sample(c("yes", "no"), 50, replace = TRUE), Z = sample(c("french","german","english"), 50, replace=TRUE) ) # Load dplyr (install first with install.packages("dplyr") if needed) library(dplyr) # Add "selection" column: mark 1 for randomly selected rows per year, 0 otherwise dat_marked <- dat %>% group_by(X) %>% # Group by year (your X column) mutate( selection = ifelse( row_number() %in% sample(n(), min(n(), 1)), # 1 = number of samples per year; replace with 53 for your project 1, 0 ) ) %>% ungroup()
For Your Actual 150k-Row Dataset
Just replace the 1 in sample(n(), min(n(), 1)) with 53. The min(n(), 53) ensures you don’t get errors if a year has fewer than 53 samples (it will select all available entries for that year):
dat_marked_large <- dat %>% group_by(X) %>% mutate( selection = ifelse( row_number() %in% sample(n(), min(n(), 53)), 1, 0 ) ) %>% ungroup()
Solution 2: Base R (No Extra Packages)
If you prefer sticking to base R, use the ave() function to handle grouping and tagging:
set.seed(123) dat$selection <- ave( 1:nrow(dat), # Pass row numbers to group by year dat$X, # Group by your year column FUN = function(x) { # Randomly select up to 53 rows per year selected_rows <- sample(x, min(length(x), 53)) # Mark selected rows as 1, others as 0 as.integer(x %in% selected_rows) } )
How to Use This
- Check selected rows: Simply filter the dataset to see which entries are marked:
# View all selected samples selected_samples <- dat_marked %>% filter(selection == 1) # Verify a specific year's selections dat_marked %>% filter(X == 2000) - Modify selected samples: If you need to edit fields for selected rows, do it directly in the original dataset (no merging required!):
dat_marked <- dat_marked %>% mutate(Y = ifelse(selection == 1, "selected_yes", Y))
Key Benefits
- No duplicate rows: We’re modifying the original dataset in-place (well, creating a single copy) instead of merging subsets.
- Memory-efficient: Works smoothly with your 150k-row dataset—no need to load huge Excel files.
- Reproducible: Using
set.seed()ensures your random selections stay consistent for future checks or collaborations.
内容的提问来源于stack exchange,提问作者C.Den

