在R语言中对DataFrame进行去重与变长转换
问题描述
合并三个数据集后得到结构混乱的DataFrame,包含唯一id字段,每个id对应一个或多个样本,原始数据结构如下:
samples <- structure(list(id = c(1029459, 1029459, 1029459, 1029459, 1030272, 1030272, 1030272, 1032157, 1032157, 1032178, 1032178, 1032219, 1032219, 1032229, 1032229, 1032494, 1032494, 1032780, 1032780 ), sample1 = c(853401, 853401, 853401, 853401, 852769, 852769, 852769, 850161, 850161, 852711, 852711, 852597, 852597, 850363, 850363, 850717, 850717, 848763, 848763), sample2 = c(853401, 853693, 853667, 853667, 852769, 853597, 853597, NA, NA, 852711, 853419, 852597, 852597, 850363, 852741, 850717, 851811, 848763, 848763), sample3 = c(NA, NA, NA, NA, NA, NA, NA, 853621, 852621, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -19L)) print(samples) #> # A tibble: 19 × 4 #> id sample1 sample2 sample3 #> <dbl> <dbl> <dbl> <dbl> #> 1 1029459 853401 853401 NA #> 2 1029459 853401 853693 NA #> 3 1029459 853401 853667 NA #> 4 1029459 853401 853667 NA #> 5 1030272 852769 852769 NA #> 6 1030272 852769 853597 NA #> 7 1030272 852769 853597 NA #> 8 1032157 850161 NA 853621 #> 9 1032157 850161 NA 852621 #> 10 1032178 852711 852711 NA #> 11 1032178 852711 853419 NA #> 12 1032219 852597 852597 NA #> 13 1032219 852597 852597 NA #> 14 1032229 850363 850363 NA #> 15 1032229 850363 852741 NA #> 16 1032494 850717 850717 NA #> 17 1032494 850717 851811 NA #> 18 1032780 848763 848763 NA #> 19 1032780 848763 848763 NA
希望将每个id对应的所有唯一样本合并到一个sample列,转换为长格式,示例输出:
id sample 1029459 853401 1029459 853693 1030272 852769 1030272 853597 1032157 850161 1032157 853621
解决方案
方法一:使用tidyverse工具链
通过pivot_longer转长格式,再去重并过滤NA值:
library(tidyverse) result <- samples %>% pivot_longer(cols = starts_with("sample"), names_to = NULL, values_to = "sample") %>% filter(!is.na(sample)) %>% distinct(id, sample) %>% arrange(id, sample) print(result)
步骤说明:
pivot_longer:将所有以sample开头的列转换为单列,丢弃原列名,值存入sample列filter(!is.na(sample)):移除样本值为NA的行distinct(id, sample):保留每个id对应的唯一样本组合arrange(id, sample):按id和sample排序,与示例输出格式一致
方法二:使用Base R
无需额外安装包,通过reshape转长格式后处理去重和NA:
# 转换为长格式 long_samples <- reshape(samples, direction = "long", varying = list(names(samples)[-1]), v.names = "sample", idvar = "id", times = NULL) # 去重、过滤NA并排序 result <- unique(long_samples[, c("id", "sample")]) result <- result[!is.na(result$sample), ] result <- result[order(result$id, result$sample), ] # 重置行名 rownames(result) <- NULL print(result)
步骤说明:
reshape:指定id为分组变量,将其余列合并为sample列unique:保留id和sample的唯一组合- 过滤NA值后按
id和sample排序,最后重置行名
内容的提问来源于stack exchange,提问作者pgcudahy
相关产品推荐
相关产品推荐

