如何简化多列组合场景下的tapply迭代代码?
Hey there! It looks like you're tired of writing repetitive tapply calls for every pair of ID columns in your data frame. Let's fix that with some dynamic code that scales automatically, no matter how many ID columns you end up adding later.
First, let's start with your sample data to make sure we're on the same page:
# Your sample data frame df <- data.frame( ID1 = c(1, 1, 1, 2, 2, 2, 3), ID2 = c(2, 3, 2, 3, 2, 3, 2), ID3 = c(2, 2, 2, 2, 1, 1, 1), Area = c(20, 30, 90, 80, 70, 67, 73) )
Approach 1: Base R (No External Packages)
This method uses base R functions to generate all ID column pairs and compute sums in a loop:
- First, we identify all columns that start with "ID" (so this works even if you add ID4, ID5, etc. later)
- Generate every possible pair of these ID columns using
utils::combn - Use
lapplyto run the sum calculation for each pair automatically
# Step 1: Extract all ID column names id_columns <- grep("^ID", names(df), value = TRUE) # Step 2: Create all 2-column combinations (returns a list of pairs) id_pairs <- utils::combn(id_columns, 2, simplify = FALSE) # Step 3: Calculate sum of Area for each pair sum_results <- lapply(id_pairs, function(pair) { tapply(df$Area, df[pair], sum) }) # Optional: Name the list elements so you know which pair each result belongs to names(sum_results) <- sapply(id_pairs, paste, collapse = " & ")
If you print sum_results, you'll get a named list where each element is the summed Area for the corresponding ID pair—exactly what you were doing manually before, but now it's automated.
Approach 2: Tidyverse (dplyr + purrr)
If you prefer working with tidy data frames instead of arrays, this method gives you clean, row-based results for each ID pair:
library(dplyr) library(purrr) # Step 1: Get ID column names id_columns <- names(df)[startsWith(names(df), "ID")] # Step 2: Generate all ID pairs id_pairs <- combn(id_columns, 2, simplify = FALSE) # Step 3: Iterate over pairs and compute grouped sums tidy_sum_results <- map(id_pairs, function(pair) { df %>% group_by(across(all_of(pair))) %>% summarise(Total_Area = sum(Area), .groups = "drop") }) # Name the list elements for clarity names(tidy_sum_results) <- map_chr(id_pairs, ~paste(.x, collapse = " & "))
Each element in tidy_sum_results is a tidy data frame with the grouped IDs and their total Area—super easy to analyze or combine later if needed.
Both approaches will scale seamlessly: if you add more ID columns (like ID4, ID5), the code will automatically generate all new pairs without any changes from you.
内容的提问来源于stack exchange,提问作者chu-js

