在R语言中生成指定长度及元素出现次数的唯一组合向量
I totally get your frustration—using permn() on a vector with duplicate elements and then calling unique() is incredibly inefficient, since you're generating tons of redundant permutations only to throw them away later. Instead, we can directly generate only the unique vectors by focusing on selecting positions for each element step by step, which avoids duplicates entirely.
The Core Idea
Instead of permuting duplicate elements, we build each unique vector by choosing positions for each number in sequence:
- First, pick 4 positions out of 14 for the number 1.
- For each set of 1's positions, pick 3 positions from the remaining 10 slots for the number 2.
- Next, pick 5 positions from the remaining 7 slots for the number 3.
- The last 2 positions automatically get the number 4.
This approach guarantees every generated vector is unique—no duplicates to clean up later.
R Code Implementation
Here's a function that implements this logic:
generate_unique_vectors <- function() { total_length <- 14 # Define the count for each number element_counts <- c("1" = 4, "2" = 3, "3" = 5, "4" = 2) # Step 1: Generate all possible positions for number 1 pos_1_list <- combn(total_length, element_counts["1"], simplify = FALSE) # Step 2: For each 1's position, generate positions for number 2 pos_1_2_list <- lapply(pos_1_list, function(pos1) { remaining_after_1 <- setdiff(1:total_length, pos1) pos_2_list <- combn(remaining_after_1, element_counts["2"], simplify = FALSE) lapply(pos_2_list, function(pos2) list(pos1 = pos1, pos2 = pos2)) }) # Flatten the nested list pos_1_2_list <- unlist(pos_1_2_list, recursive = FALSE) # Step 3: For each 1+2 positions, generate positions for 3 and build the vector final_vectors <- lapply(pos_1_2_list, function(positions) { remaining_after_1_2 <- setdiff(1:total_length, c(positions$pos1, positions$pos2)) pos_3_list <- combn(remaining_after_1_2, element_counts["3"], simplify = FALSE) lapply(pos_3_list, function(pos3) { vec <- integer(total_length) vec[positions$pos1] <- 1 vec[positions$pos2] <- 2 vec[pos3] <- 3 # Fill remaining slots with 4 vec[setdiff(1:total_length, c(positions$pos1, positions$pos2, pos3))] <- 4 vec }) }) # Flatten to get a list of unique vectors final_vectors <- unlist(final_vectors, recursive = FALSE) return(final_vectors) } # Generate all unique vectors unique_vecs <- generate_unique_vectors() # Check the first few results head(unique_vecs)
Verify the Result
You can confirm that we're generating exactly the right number of unique vectors by comparing the theoretical count to the length of the output:
# Calculate theoretical total (matches your formula) theoretical_total <- choose(14,4) * choose(10,3) * choose(7,5) * choose(2,2) cat("Theoretical total:", theoretical_total, "\n") cat("Actual generated count:", length(unique_vecs), "\n")
These two numbers will be identical, proving we have no duplicates and no missing vectors.
Why This Is Better
This method is far more efficient than generating permutations and deduplicating:
- We skip generating redundant permutations entirely, saving memory and computation time.
- For cases with many duplicate elements, the performance gap becomes massive—you won't waste resources on permutations that are just copies of each other.
内容的提问来源于stack exchange,提问作者Vassilis Chasiotis

