You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R原DataFrame中标记按年随机抽取的样本?

Hey there! Let's work through this problem—since you're handling large biological archival data, we need a clean, memory-efficient way to mark your selected rows directly in the original DataFrame, no more messy duplicates or Excel workarounds.

Core Idea

Instead of extracting subsets and merging back (which causes duplicate issues), we'll add a marker column directly to your original dataset by grouping samples by year, then randomly selecting the required number of entries per group and tagging them.


Solution 1: Using dplyr (Clean & Readable)

This is great for code clarity, and it’s efficient even for large datasets like yours.

First, let’s adapt your sample code to mark selected rows:

# Set seed for reproducible results (critical for research!)
set.seed(123)

# Your sample dataset
dat <- data.frame(
  X = sample(2000:2016, 50, replace=TRUE),
  Y = sample(c("yes", "no"), 50, replace = TRUE),
  Z = sample(c("french","german","english"), 50, replace=TRUE)
)

# Load dplyr (install first with install.packages("dplyr") if needed)
library(dplyr)

# Add "selection" column: mark 1 for randomly selected rows per year, 0 otherwise
dat_marked <- dat %>%
  group_by(X) %>%  # Group by year (your X column)
  mutate(
    selection = ifelse(
      row_number() %in% sample(n(), min(n(), 1)),  # 1 = number of samples per year; replace with 53 for your project
      1,
      0
    )
  ) %>%
  ungroup()

For Your Actual 150k-Row Dataset

Just replace the 1 in sample(n(), min(n(), 1)) with 53. The min(n(), 53) ensures you don’t get errors if a year has fewer than 53 samples (it will select all available entries for that year):

dat_marked_large <- dat %>%
  group_by(X) %>%
  mutate(
    selection = ifelse(
      row_number() %in% sample(n(), min(n(), 53)),
      1,
      0
    )
  ) %>%
  ungroup()

Solution 2: Base R (No Extra Packages)

If you prefer sticking to base R, use the ave() function to handle grouping and tagging:

set.seed(123)

dat$selection <- ave(
  1:nrow(dat),  # Pass row numbers to group by year
  dat$X,        # Group by your year column
  FUN = function(x) {
    # Randomly select up to 53 rows per year
    selected_rows <- sample(x, min(length(x), 53))
    # Mark selected rows as 1, others as 0
    as.integer(x %in% selected_rows)
  }
)

How to Use This

  • Check selected rows: Simply filter the dataset to see which entries are marked:
    # View all selected samples
    selected_samples <- dat_marked %>% filter(selection == 1)
    
    # Verify a specific year's selections
    dat_marked %>% filter(X == 2000)
    
  • Modify selected samples: If you need to edit fields for selected rows, do it directly in the original dataset (no merging required!):
    dat_marked <- dat_marked %>%
      mutate(Y = ifelse(selection == 1, "selected_yes", Y))
    

Key Benefits

  1. No duplicate rows: We’re modifying the original dataset in-place (well, creating a single copy) instead of merging subsets.
  2. Memory-efficient: Works smoothly with your 150k-row dataset—no need to load huge Excel files.
  3. Reproducible: Using set.seed() ensures your random selections stay consistent for future checks or collaborations.

内容的提问来源于stack exchange,提问作者C.Den

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:52:23