R语言:如何按x1的键变量分组合并df1与df2且无重复行?
Sure thing! You can absolutely combine df1 and df2 grouped by the key variable x1 in the way you described—getting a df1 + df2 style result without duplicate rows, no need to worry about relationships between new columns. Let me show you two straightforward, practical approaches using common R packages:
1. Using dplyr + purrr
This approach splits both dataframes by x1, combines corresponding groups, removes duplicates, then binds everything back together. First, let’s use sample data to illustrate:
# Sample input data df1 <- data.frame(x1 = c("A", "A", "B"), val1 = c(1, 2, 3)) df2 <- data.frame(x1 = c("A", "B", "B"), val2 = c(10, 20, 30)) # Load required packages library(dplyr) library(purrr) # Split each dataframe into groups based on x1 split_df1 <- df1 %>% group_split(x1, .keep = TRUE) split_df2 <- df2 %>% group_split(x1, .keep = TRUE) # Combine matching groups, remove duplicates, then bind all groups combined_groups <- map2(split_df1, split_df2, ~bind_rows(.x, .y) %>% distinct()) final_df <- bind_rows(combined_groups)
The distinct() step ensures no duplicate rows within each group, and the result will include all rows from both dataframes, organized by x1.
2. Using data.table (more efficient for large datasets)
If you’re working with big data, data.table offers a concise and fast solution:
# Sample input data (same as above) df1 <- data.frame(x1 = c("A", "A", "B"), val1 = c(1, 2, 3)) df2 <- data.frame(x1 = c("A", "B", "B"), val2 = c(10, 20, 30)) # Load package and convert to data.table library(data.table) setDT(df1) setDT(df2) # Combine dataframes, then deduplicate within each x1 group final_df <- rbind(df1, df2)[, .SD %>% distinct(), by = x1]
This one-liner first stacks the two dataframes, then removes duplicate rows within each x1 group—exactly what you need.
Example Output
For the sample data above, final_df will look like this:
x1 val1 val2 1: A 1 NA 2: A 2 NA 3: A NA 10 4: B 3 NA 5: B NA 20 6: B NA 30
If there were identical rows across df1 and df2 (same values in all columns), distinct() would automatically drop the duplicates.
内容的提问来源于stack exchange,提问作者Fabian H

