如何在Base R中无重复抽取随机样本,兼顾统计可靠性与结果可复现性?
Great question! Let's break this down and tackle both of your questions, plus weigh in on the approach you're considering.
Core Recap of Your Needs
You want to:
- Generate non-repeating random samples (no duplicate individuals within a single sample, and no duplicate full sample sets across multiple draws)
- Use only Base R (no external packages)
- Have fully reproducible results
- Ensure the method is statistically reliable
First: Why Your Seed-Based Approach Doesn’t Work
Generating 1000 random seeds and looping through them to draw samples won’t guarantee unique sample sets. Different seeds can still produce identical sampling outcomes (even if the probability is low, it’s not zero), and there’s no built-in check to filter out duplicates. So this method can’t reliably meet your requirement of avoiding duplicate sample sets.
Solution 1: For Small Populations (Feasible to List All Possible Samples)
If your population size is small enough that you can generate all possible non-replacement samples, you can directly sample from this full set of combinations to get unique sample sets. This is statistically sound and fully reproducible.
Example code:
# Set seed for reproducibility set.seed(1) # Define your population and sampling parameters population <- 1:100 # 100 individuals total sample_size <- 10 # Draw 10 individuals per sample num_unique_samples <- 5 # Want 5 unique sample sets # Generate ALL possible non-replacement sample index combinations all_possible_samples <- combn(length(population), sample_size, simplify = FALSE) # Randomly select num_unique_samples combinations (without replacement) selected_indices <- sample(length(all_possible_samples), num_unique_samples, replace = FALSE) # Extract the actual sample sets from the population unique_samples <- lapply(selected_indices, function(i) population[all_possible_samples[[i]]]) # Check the results unique_samples
Note: This only works for small populations. If your population is large (e.g., 1000 individuals) and sample size is meaningful, the number of possible combinations becomes astronomically large—generating all of them is impossible due to memory constraints.
Solution 2: For Large Populations (Practical & Statistically Reliable)
For larger populations, we use a "draw, check, and keep" approach. Since the probability of drawing an identical sample set is extremely low with large populations, this loop runs efficiently, and we explicitly filter out duplicates to guarantee uniqueness.
Example code:
# Set seed for reproducibility set.seed(1) # Define parameters population <- 1:1000 # Large population of 1000 individuals sample_size <- 100 # Draw 100 individuals per sample num_unique_samples <- 1000 # Need 1000 unique sample sets # Initialize a list to store our unique samples unique_samples <- list() # Loop until we have enough unique samples while (length(unique_samples) < num_unique_samples) { # Draw a single non-replacement sample current_sample <- sample(population, sample_size, replace = FALSE) # Sort the sample to compare consistency (since [1,2] and [2,1] are the same sample set) sorted_current <- sort(current_sample) # Check if this sample set already exists in our list is_duplicate <- any(sapply(unique_samples, function(x) identical(sort(x), sorted_current))) # If it's not a duplicate, add it to the list if (!is_duplicate) { unique_samples[[length(unique_samples) + 1]] <- current_sample } } # Verify all samples are unique all_unique <- !duplicated(lapply(unique_samples, sort)) all(all_unique) # Should return TRUE
Why This Is Statistically Reliable:
- Each initial draw uses Base R's
sample()withreplace = FALSE, which is a standard, statistically valid simple random sample without replacement. - We only filter out duplicate sample sets—this doesn’t introduce bias, since we’re just excluding rare repeat outcomes, not altering the randomness of the draws themselves.
Key Takeaways
- Reproducibility: Always start with
set.seed()to ensure your results can be replicated exactly. - Within-Sample Uniqueness: Use
sample(population, size, replace = FALSE)to guarantee no duplicate individuals in a single sample. - Across-Sample Uniqueness: Use combination sampling for small populations, or the draw-check-keep loop for large populations.
内容的提问来源于stack exchange,提问作者user6472523

