R中按分组rbind列表列嵌套tibble的惯用实现方法
问题描述
我有一个tibble对象,其中包含一个存储tibble的列表列,各嵌套tibble的列结构兼容。我希望先按指定列分组后,对组内的嵌套tibble执行rbind行合并操作。以下是简化示例,本次我需要按tpm列分组:
library(tidyverse) df_ex <- structure(list( tpm = c(3, 3, 5, 5), strand = c("negative", "positive", "negative", "positive"), sites = list( structure(list(chr = c("1", "1"), pos = c(30214L, 31109L), cov = c(7L, 14L), strand = c("-", "-")), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame")), structure(list(chr = c("1", "1"), pos = c(14362L, 14406L), cov = c(130L, 5490L), strand = c("+", "+")), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame")), structure(list(chr = c("1", "1"), pos = c(96976L, 98430L), cov = c(185L,3L), strand = c("-", "-")), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame")), structure(list(chr = c("1", "1"), pos = c(14358L, 14406L), cov = c(24L, 5246L), strand = c("+", "+")), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame")))), row.names = c(NA, -4L), class = c("tbl_df", "tbl", "data.frame")) df_ex ## A tibble: 4 × 3 # tpm strand sites # <dbl> <chr> <list> # 1 3 negative <tibble [2 × 4]> # 2 3 positive <tibble [2 × 4]> # 3 5 negative <tibble [2 × 4]> # 4 5 positive <tibble [2 × 4]>
失败尝试
写法1:使用transmute
df_ex %>% group_by(tpm) %>% transmute(sites=do.call(rbind, sites))
运行后抛出错误:
Error in `transmute()`: ! Problem while computing `sites = do.call(rbind, sites)`. ✖ `sites` must be size 2 or 1, not 4. ℹ The error occurred in group 1: tpm = 3. Run `rlang::last_error()` to see where the error occurred.
写法2:使用summarize
df_ex %>% group_by(tpm) %>% summarize(sites=do.call(rbind, sites), .groups='drop')
该写法会自动展开嵌套的tibble结构,得到扁平化的8行结果,不符合预期:
# A tibble: 8 × 2 tpm sites$chr $pos $cov $strand <dbl> <chr> <int> <int> <chr> 1 3 1 30214 7 - 2 3 1 31109 14 - 3 3 1 14362 130 + 4 3 1 14406 5490 + 5 5 1 96976 185 - 6 5 1 98430 3 - 7 5 1 14358 24 + 8 5 1 14406 5246 +
预期结果
按tpm分组后,sites仍保持列表列格式,每个元素存储对应组行合并后的tibble,效果如下:
## A tibble: 2 × 2 # tpm sites # <dbl> <list> # 1 3 <tibble [4 × 4]> # 2 5 <tibble [4 × 4]>
解决方案
符合tidyverse规范的实现方式是在summarise中,将组内行合并的结果用list()包裹,强制保留列表列结构。推荐使用dplyr::bind_rows替代原生rbind,对tibble的兼容性更好:
df_ex %>% group_by(tpm) %>% summarise( sites = list(bind_rows(sites)), .groups = "drop" )
运行输出与预期完全一致:
# A tibble: 2 × 2 tpm sites <dbl> <list> 1 3 <tibble [4 × 4]> 2 5 <tibble [4 × 4]>
原写法错误原因
transmute报错:transmute要求输出长度与分组输入的行数一致,tpm=3的组内原有2行数据,但do.call(rbind, sites)返回了4行合并结果,长度不匹配触发校验错误。summarise返回扁平化结果:当summarise接收到多行数据框作为汇总结果时,会自动将嵌套结构打平为普通列。在外层包裹list()后,相当于每个分组返回1个列表元素,tidyverse就会保留列表列格式,不会自动展开。
如果偏好purrr生态的函数,也可以用purrr::list_rbind实现行合并,效果完全相同:
df_ex %>% group_by(tpm) %>% summarise( sites = list(list_rbind(sites)), .groups = "drop" )
内容的提问来源于stack exchange,提问作者merv
相关产品推荐
相关产品推荐

