You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中保留DataFrame里col1和col2重复的所有行

筛选DataFrame中指定列组合重复的所有行

你需要筛选的是col1和col2组合出现多次的所有行,以下是两种常用工具的实现方式:

Pandas(Python)

利用duplicated()函数的keep=False参数,该参数会标记所有重复的行(包括首次出现的),以此作为筛选条件即可:

import pandas as pd

# 构造示例数据
df = pd.DataFrame({
    'col1': ['tn1', 'tn1', 'tn2', 'tn3'],
    'col2': ['a', 'a', 'd', 'a'],
    'col3': ['b', 'c', 'b', 'b']
})

# 筛选col1&col2组合重复的所有行
filtered_df = df[df.duplicated(subset=['col1', 'col2'], keep=False)]
print(filtered_df)

输出结果:

col1 col2 col3
0  tn1    a    b
1  tn1    a    c

dplyr(R)

通过分组后筛选组内行数大于1的记录,保留组内所有行:

library(dplyr)

# 构造示例数据
df <- data.frame(
  col1 = c("tn1", "tn1", "tn2", "tn3"),
  col2 = c("a", "a", "d", "a"),
  col3 = c("b", "c", "b", "b")
)

# 筛选col1&col2组合重复的所有行
filtered_df <- df %>%
  group_by(col1, col2) %>%
  filter(n() > 1) %>%
  ungroup()

print(filtered_df)

输出结果:

# A tibble: 2 × 3
  col1  col2  col3 
  <chr> <chr> <chr>
1 tn1   a     b    
2 tn1   a     c    

为什么之前的方法不适用?

  • unique()/distinct():作用是去重,保留唯一的行,和你需要保留重复行的需求完全相反;
  • anti_join():用于筛选在另一个表中不存在的行,无法实现筛选指定列组合重复行的功能。

内容的提问来源于stack exchange,提问作者Pame

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 17:30:59