如何在R中清理数据框重复列?保留值不同组的单列
处理含重复根名称列的Data Frame:保留值唯一的列
问题描述
有一个包含重复根名称列的Data Frame,R会自动在重复列名后追加数字保证唯一性。需要对每个根名称对应的列组进行处理:
- 若列内值完全相同,仅保留一列
- 若存在值不同的列,每组保留所有值唯一的列
样本数据
df <- data.frame( ID = rep(1:4, each = 1), CMW = rep(c(10, 20, 30, 30), each = 1), D_D = c(rep(100, 3), 200), D_D = c(rep(100, 3), 200), D_D = c(rep(100, 1), 200), Eref = rep(4:4, each = 1), Eref = rep(4:4, each = 1), Eref = rep(1:4, each = 1), Eref = rep(1:4, each = 1) )
原数据输出:
ID CMW D_D D_D.1 D_D.2 Eref Eref.1 Eref.2 Eref.3 1 1 10 100 100 100 4 4 1 1 2 2 20 100 100 200 4 4 2 2 3 3 30 100 100 100 4 4 3 3 4 4 30 200 200 200 4 4 4 4
解决方案
方法1:使用dplyr和purrr
library(dplyr) library(purrr) # 提取列的根名称(去掉后缀的数字) col_roots <- sub("\\.\\d+$", "", names(df)) # 按根名称分组,保留每组内值唯一的列 result_cols <- split(names(df), col_roots) %>% map(function(cols) { # 取出当前组的所有列 group_df <- df[, cols, drop = FALSE] # 转置后去重,保留内容唯一的列名 group_df %>% t() %>% as.data.frame() %>% distinct() %>% rownames() }) %>% unlist() # 生成结果Data Frame df_result <- df[, result_cols]
方法2:Base R实现
如果不想加载额外包,可以用Base R完成:
# 提取列的根名称 col_roots <- sub("\\.\\d+$", "", names(df)) # 按根名称分组列名 col_groups <- split(names(df), col_roots) result_cols <- c() for (group in col_groups) { # 取出当前组的列数据 group_data <- df[, group] # 将每列转为字符串,判断内容是否重复 col_hashes <- apply(group_data, 2, function(x) paste(x, collapse = ",")) # 保留不重复的列 unique_cols <- group[!duplicated(col_hashes)] result_cols <- c(result_cols, unique_cols) } # 生成结果Data Frame df_result <- df[, result_cols]
结果验证
处理后的结果:
ID CMW D_D D_D.2 Eref Eref.2 1 1 10 100 100 4 1 2 2 20 100 200 4 2 3 3 30 100 100 4 3 4 4 30 200 200 4 4
结果符合需求:同根名称的列中,值相同的仅保留一列,值不同的列全部保留。
内容的提问来源于stack exchange,提问作者Raquel Feltrin
相关产品推荐
相关产品推荐

