You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:如何按x1的键变量分组合并df1与df2且无重复行?

Sure thing! You can absolutely combine df1 and df2 grouped by the key variable x1 in the way you described—getting a df1 + df2 style result without duplicate rows, no need to worry about relationships between new columns. Let me show you two straightforward, practical approaches using common R packages:

1. Using dplyr + purrr

This approach splits both dataframes by x1, combines corresponding groups, removes duplicates, then binds everything back together. First, let’s use sample data to illustrate:

# Sample input data
df1 <- data.frame(x1 = c("A", "A", "B"), val1 = c(1, 2, 3))
df2 <- data.frame(x1 = c("A", "B", "B"), val2 = c(10, 20, 30))

# Load required packages
library(dplyr)
library(purrr)

# Split each dataframe into groups based on x1
split_df1 <- df1 %>% group_split(x1, .keep = TRUE)
split_df2 <- df2 %>% group_split(x1, .keep = TRUE)

# Combine matching groups, remove duplicates, then bind all groups
combined_groups <- map2(split_df1, split_df2, ~bind_rows(.x, .y) %>% distinct())
final_df <- bind_rows(combined_groups)

The distinct() step ensures no duplicate rows within each group, and the result will include all rows from both dataframes, organized by x1.

2. Using data.table (more efficient for large datasets)

If you’re working with big data, data.table offers a concise and fast solution:

# Sample input data (same as above)
df1 <- data.frame(x1 = c("A", "A", "B"), val1 = c(1, 2, 3))
df2 <- data.frame(x1 = c("A", "B", "B"), val2 = c(10, 20, 30))

# Load package and convert to data.table
library(data.table)
setDT(df1)
setDT(df2)

# Combine dataframes, then deduplicate within each x1 group
final_df <- rbind(df1, df2)[, .SD %>% distinct(), by = x1]

This one-liner first stacks the two dataframes, then removes duplicate rows within each x1 group—exactly what you need.

Example Output

For the sample data above, final_df will look like this:

x1 val1 val2
1:  A    1   NA
2:  A    2   NA
3:  A   NA   10
4:  B    3   NA
5:  B   NA   20
6:  B   NA   30

If there were identical rows across df1 and df2 (same values in all columns), distinct() would automatically drop the duplicates.

内容的提问来源于stack exchange,提问作者Fabian H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:56:58