You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中合并数据框(tibble)中的列对?

基因型数据两两列合并解决方案

我需要处理基因型数据:原始数据框中每个样本对应两列(分别代表两个等位基因),要将每两列用|分隔合并为一列,最终得到n/2列的结果。之前尝试tidyr::unite会把所有列合并成一列,不符合需求。

原始数据

c1.1  c1.2  c2.1  c2.2  c3.1  c3.2  c4.1  c4.2  c5.1  c5.2
 1     0     0     0     0     0     0     0     0     0     0
 2     0     0     0     0     0     0     0     0     0     1
 3     0     0     0     0     0     0     0     0     0     0
 4     0     1     0     1     0     1     1     1     1     1
 5     0     0     0     0     0     0     0     0     0     0
 6     0     0     0     0     0     0     0     0     0     0
 7     0     0     0     0     0     0     0     0     0     0
 8     0     1     0     1     0     1     0     0     0     0
 9     0     0     0     0     0     0     0     0     0     0
10     0     0     0     0     0     0     0     0     0     0

目标结果

c1      c2      c3      c4      c5
1     0|0     0|0     0|0     0|0     0|0
2     0|0     0|0     0|0     0|0     0|1
3     0|0     0|0     0|0     0|0     0|0
4     0|1     0|1     0|1     1|1     1|1
5     0|0     0|0     0|0     0|0     0|0
6     0|0     0|0     0|0     0|0     0|0
7     0|0     0|0     0|0     0|0     0|0
8     0|1     0|1     0|1     0|0     0|0
9     0|0     0|0     0|0     0|0     0|0
10    0|0     0|0     0|0     0|0     0|0

示例数据生成代码

df <- matrix(c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
               0, 0, 0, 1, 0, 0, 0, 1, 0, 0,
               0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
               0, 0, 0, 1, 0, 0, 0, 1, 0, 0,
               0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
               0, 0, 0, 1, 0, 0, 0, 1, 0, 0,
               0, 0, 0, 1, 0, 0, 0, 0, 0, 0,
               0, 0, 0, 1, 0, 0, 0, 0, 0, 0,
               0, 0, 0, 1, 0, 0, 0, 0, 0, 0,
               0, 1, 0, 1, 0, 0, 0, 0, 0, 0),
             ncol = 10)

df <- dplyr::as_tibble(df)

samples <- c("c1","c2","c3","c4","c5")

names(df) <- paste0(rep(samples, each = 2), c(".1",".2"))

可行解法

方法1:dplyr + tidyr 重塑数据

通过行列转换实现分组合并:

library(dplyr)
library(tidyr)

result <- df %>%
  # 添加行标识符,避免合并时丢失行顺序
  mutate(row_id = row_number()) %>%
  # 宽表转长表,拆分列名为分组和后缀
  pivot_longer(-row_id, names_to = c("group", "suffix"), names_sep = "\\.") %>%
  # 长表转宽表,同一分组的两个值用|合并
  pivot_wider(
    names_from = group,
    values_from = value,
    values_fn = \(x) paste(x, collapse = "|")
  ) %>%
  # 移除行标识符
  select(-row_id)

方法2:purrr 批量处理每组列

利用map_dfc循环处理每个样本对应的两列:

library(purrr)
library(dplyr)

result <- map_dfc(samples, function(sample) {
  df %>%
    select(starts_with(sample)) %>%
    unite(col = sample, everything(), sep = "|")
})

方法3:Base R 原生实现

无需加载额外包,直接循环处理:

result <- data.frame()
for (s in samples) {
  # 匹配当前样本的两列
  col_indices <- grep(paste0("^", s, "\\."), names(df))
  # 逐行合并两列值
  merged_col <- apply(df[, col_indices], 1, \(x) paste(x, collapse = "|"))
  result <- cbind(result, merged_col)
}
# 设置列名并转为tibble
names(result) <- samples
result <- dplyr::as_tibble(result)

内容的提问来源于stack exchange,提问作者TTuovinen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 18:48:10