如何在R中计算两个DataFrame行间倍数变化并生成新DataFrame?
问题描述
我有两个行数不同的DataFrame,需要将第一个DataFrame的每一行与第二个DataFrame的每一行两两配对比较,计算Count值的倍数增减情况,最终生成包含该结果的新DataFrame。
两个DataFrame的结构如下:
# 第一个DataFrame structure(list(Label = c("Gene 1", "Gene 2", "Gene 3", "Gene 4", "Gene 5", "Gene 6", "Gene 7", "Gene 8", "Gene 9", "Gene 10", "Gene 11", "Gene 12", "Gene 13", "Gene 14", "Gene 15", "Gene 16", "Gene 17", "Gene 18", "Gene 19", "Gene 20", "Gene 21", "Gene 22", "Gene 23", "Gene 24", "Gene 25", "Gene 26", "Gene 27", "Gene 28", "Gene 29", "Gene 30"), Count = c(1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -30L)) # 第二个DataFrame structure(list(Label = c("Control1", "Control2", "Control3", "Control4", "Control5", "Control6", "Control7", "Control8", "Control9", "Control10", "Control11", "Control12", "Control13", "Control14", "Control15", "Control16", "Control17", "Control18", "Control19", "Control20", "Control21", "Control22", "Control23", "Control24" ), Count = c(1800, 1400, 1110, 1900, 2500, 2900, 2100, 900, 5000, 2300, 700, 1400, 3400, 2310, 3322, 2200, 4400, 2100, 1000, 6700, 4300, 2120, 4800, 4300)), class = c("tbl_df", "tbl", "data.frame" ), row.names = c(NA, -24L))
期望输出为包含基因标签、对照标签、Count倍数关系的表格,请问有没有现成函数可以实现该需求,无需手动编写循环?
解决方案
可以借助tidyverse工具包的cross_join()函数快速生成全组合并计算倍数,完全不需要循环。
步骤1:加载工具包并读入数据
假设两个DataFrame分别命名为gene_df和control_df:
library(tidyverse) # 读入第一个DataFrame gene_df <- structure(list(Label = c("Gene 1", "Gene 2", ...), Count = c(1500, 1600, ...)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -30L)) # 读入第二个DataFrame control_df <- structure(list(Label = c("Control1", "Control2", ...), Count = c(1800, 1400, ...)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -24L))
步骤2:生成全组合并计算倍数
result_df <- cross_join(gene_df, control_df, suffix = c("_gene", "_control")) %>% mutate(Count_ratio = Count_gene / Count_control) %>% select(Gene_Label = Label_gene, Control_Label = Label_control, Count_gene, Count_control, Count_ratio)
cross_join()自动生成两个DataFrame所有行的笛卡尔积配对;suffix参数用来区分两个DataFrame中重名的列(这里是Label和Count);mutate()计算倍数比值,若需要反向倍数可改为Count_control / Count_gene;select()调整列名与顺序,让结果更直观。
无需额外包的base R实现
如果不想加载tidyverse,可以用expand.grid()生成组合后合并计算:
# 生成所有行的索引组合 all_pairs <- expand.grid(gene_row = 1:nrow(gene_df), control_row = 1:nrow(control_df)) # 合并数据并计算倍数 result_df <- data.frame( Gene_Label = gene_df$Label[all_pairs$gene_row], Control_Label = control_df$Label[all_pairs$control_row], Count_gene = gene_df$Count[all_pairs$gene_row], Count_control = control_df$Count[all_pairs$control_row], Count_ratio = gene_df$Count[all_pairs$gene_row] / control_df$Count[all_pairs$control_row] )
两种方法都能高效生成目标结果,比手动循环更简洁且性能更优。
内容的提问来源于stack exchange,提问作者Abhi
相关产品推荐
相关产品推荐

