You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Base R中无重复抽取随机样本,兼顾统计可靠性与结果可复现性?

Great question! Let's break this down and tackle both of your questions, plus weigh in on the approach you're considering.

Core Recap of Your Needs

You want to:

  • Generate non-repeating random samples (no duplicate individuals within a single sample, and no duplicate full sample sets across multiple draws)
  • Use only Base R (no external packages)
  • Have fully reproducible results
  • Ensure the method is statistically reliable

First: Why Your Seed-Based Approach Doesn’t Work

Generating 1000 random seeds and looping through them to draw samples won’t guarantee unique sample sets. Different seeds can still produce identical sampling outcomes (even if the probability is low, it’s not zero), and there’s no built-in check to filter out duplicates. So this method can’t reliably meet your requirement of avoiding duplicate sample sets.


Solution 1: For Small Populations (Feasible to List All Possible Samples)

If your population size is small enough that you can generate all possible non-replacement samples, you can directly sample from this full set of combinations to get unique sample sets. This is statistically sound and fully reproducible.

Example code:

# Set seed for reproducibility
set.seed(1)

# Define your population and sampling parameters
population <- 1:100  # 100 individuals total
sample_size <- 10    # Draw 10 individuals per sample
num_unique_samples <- 5  # Want 5 unique sample sets

# Generate ALL possible non-replacement sample index combinations
all_possible_samples <- combn(length(population), sample_size, simplify = FALSE)

# Randomly select num_unique_samples combinations (without replacement)
selected_indices <- sample(length(all_possible_samples), num_unique_samples, replace = FALSE)

# Extract the actual sample sets from the population
unique_samples <- lapply(selected_indices, function(i) population[all_possible_samples[[i]]])

# Check the results
unique_samples

Note: This only works for small populations. If your population is large (e.g., 1000 individuals) and sample size is meaningful, the number of possible combinations becomes astronomically large—generating all of them is impossible due to memory constraints.


Solution 2: For Large Populations (Practical & Statistically Reliable)

For larger populations, we use a "draw, check, and keep" approach. Since the probability of drawing an identical sample set is extremely low with large populations, this loop runs efficiently, and we explicitly filter out duplicates to guarantee uniqueness.

Example code:

# Set seed for reproducibility
set.seed(1)

# Define parameters
population <- 1:1000  # Large population of 1000 individuals
sample_size <- 100    # Draw 100 individuals per sample
num_unique_samples <- 1000  # Need 1000 unique sample sets

# Initialize a list to store our unique samples
unique_samples <- list()

# Loop until we have enough unique samples
while (length(unique_samples) < num_unique_samples) {
  # Draw a single non-replacement sample
  current_sample <- sample(population, sample_size, replace = FALSE)
  # Sort the sample to compare consistency (since [1,2] and [2,1] are the same sample set)
  sorted_current <- sort(current_sample)
  
  # Check if this sample set already exists in our list
  is_duplicate <- any(sapply(unique_samples, function(x) identical(sort(x), sorted_current)))
  
  # If it's not a duplicate, add it to the list
  if (!is_duplicate) {
    unique_samples[[length(unique_samples) + 1]] <- current_sample
  }
}

# Verify all samples are unique
all_unique <- !duplicated(lapply(unique_samples, sort))
all(all_unique)  # Should return TRUE

Why This Is Statistically Reliable:

  • Each initial draw uses Base R's sample() with replace = FALSE, which is a standard, statistically valid simple random sample without replacement.
  • We only filter out duplicate sample sets—this doesn’t introduce bias, since we’re just excluding rare repeat outcomes, not altering the randomness of the draws themselves.

Key Takeaways

  • Reproducibility: Always start with set.seed() to ensure your results can be replicated exactly.
  • Within-Sample Uniqueness: Use sample(population, size, replace = FALSE) to guarantee no duplicate individuals in a single sample.
  • Across-Sample Uniqueness: Use combination sampling for small populations, or the draw-check-keep loop for large populations.

内容的提问来源于stack exchange,提问作者user6472523

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:25:29