在R中按group分组统计每组内的unique gene数量
问题:统计每个唯一group对应的唯一gene数量
需求:根据提供的example数据集,统计每个group下的唯一gene数量,预期输出为:
- "Amino acid metabolism" 和 "Carbohydrate metabolism" 组的
total值为2 - "Translation" 和 "Membrane transport" 组的
total值为1
用户尝试的两段代码均未得到预期结果:
尝试代码1
library(tidyverse) example %>% group_by(group,gene) %>% mutate(total = n()) %>% unique() -> count.viz
尝试代码2
example %>% select(gene) %>% group_by(group) %>% mutate(total = n()) %>% unique() -> count.viz
数据集的dput结构:
> dput(example) structure(list(KO = c("K00639", "K12308", "K01881", "K16785", "K03639", "K00174"), Project = c("Ga0598246", "Ga0598246", "Ga0598246", "Ga0598246", "Ga0598246", "Ga0598246"), count = c(0, 0, 0, 0, 0, 0.0170424288146159), gene = c("kbl, GCAT", "bgaB, lacA", "PARS, proS", "ecfT", "moaA, CNX2", "korA, oorA, oforA"), group = c("Amino acid metabolism", "Carbohydrate metabolism", "Translation", "Membrane transport", "Amino acid metabolism", "Carbohydrate metabolism")), class = c("grouped_df", "tbl_df", "tbl", "data.frame"), row.names = c(NA, -6L), groups = structure(list( Project = "Ga0598246", .rows = structure(list(1:6), ptype = integer(0), class = c("vctrs_list_of", "vctrs_vctr", "list"))), class = c("tbl_df", "tbl", "data.frame" ), row.names = c(NA, -1L), .drop = TRUE))
错误代码分析
- 代码1:按
group和gene联合分组后,n()统计的是每个(group, gene)组合的出现次数,而非每个group下的唯一gene总数,最终结果会保留所有gene行,不符合需求。 - 代码2:先执行
select(gene)丢失了group列,后续group_by(group)会直接报错;即使保留group列,n()统计的是该group下的总行数,而非唯一gene的数量。
正确实现方法
以下两种tidyverse方法均可得到预期结果:
方法1:先去重再分组计数
先保留group和gene的唯一组合,再按group统计数量:
library(tidyverse) count.viz <- example %>% distinct(group, gene) %>% # 保留唯一的group-gene组合 group_by(group) %>% summarize(total = n())
方法2:直接使用n_distinct()统计唯一值
在分组后,用n_distinct(gene)直接计算每个group下的唯一gene数量:
count.viz <- example %>% group_by(group) %>% summarize(total = n_distinct(gene))
两种方法的输出结果一致:
# A tibble: 4 × 2 group total <chr> <int> 1 Amino acid metabolism 2 2 Carbohydrate metabolism 2 3 Membrane transport 1 4 Translation 1
内容的提问来源于stack exchange,提问作者Geomicro
相关产品推荐
相关产品推荐

