You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中识别、移除重复列并按组统计重复列信息

识别并分组R语言数据框中的重复列

你当前的代码能完成重复列移除,但没法把重复列按组归类。下面提供两种方案,既能输出你要的分组结果,也能保留去重后的数据集:

方案一:Base R 原生实现

# 0. 创建模拟数据集
my_df <- data.frame(
  x1 = c(1:5),
  x2 = letters[1:5],
  x3 = seq(10,50,by=10),
  x4 = c(1:5),
  x5 = c(1:5),
  x6 = LETTERS[1:5],
  x7 = letters[1:5],
  x8 = c(1:5)
)

# 1. 给每列生成唯一特征串,用来判断是否重复
col_signatures <- sapply(my_df, function(col) paste(col, collapse = "|"))

# 2. 按特征串分组,提取每组的列名,过滤掉无重复的单列表
duplicate_groups <- split(names(col_signatures), col_signatures)
duplicate_groups <- duplicate_groups[sapply(duplicate_groups, length) > 1]

# 查看分组结果
duplicate_groups
# 输出结果:
# $`1|2|3|4|5`
# [1] "x1" "x4" "x5" "x8"
# 
# $`a|b|c|d|e`
# [1] "x2" "x7"

# 3. 移除重复列(保留每组第一列)
unique_df <- my_df[!duplicated(col_signatures)]

方案二:tidyverse 工具链实现

如果平时常用dplyr这类包,用下面的代码更顺手:

library(dplyr)
library(tidyr)

# 模拟数据集同上

# 1. 转置数据后分组,提取重复列组
duplicate_groups <- my_df %>%
  t() %>%
  as.data.frame() %>%
  rownames_to_column("col_name") %>%
  group_by(across(-col_name)) %>%
  summarise(duplicate_cols = list(col_name), .groups = "drop") %>%
  filter(length(duplicate_cols) > 1) %>%
  pull(duplicate_cols)

# 查看分组结果
duplicate_groups
# 输出结果:
# [[1]]
# [1] "x1" "x4" "x5" "x8"
# 
# [[2]]
# [1] "x2" "x7"

# 2. 移除重复列
unique_df <- my_df %>% select(unique(col_signatures))

两种方案都能生成你需要的分组格式,同时完成去重。方案一不用额外装包,适合轻量场景;方案二代码更直观,符合tidyverse的操作逻辑。

内容的提问来源于stack exchange,提问作者aspire57

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 23:37:30