按站点分组计算各处理组与对照组的Bray-Curtis相异度
解决方案:按站点分组计算处理组与对照组的Bray-Curtis相异度
可以通过分组处理+距离矩阵提取的方式高效实现需求,避免手动拆分站点的繁琐操作。以下是基于dplyr、purrr和vegan包的解决方案:
步骤1:加载所需包与模拟数据
library(dplyr) library(purrr) library(vegan) # 模拟数据 Site = c("A", "A", "A", "B", "B", "B", "C", "C", "C") Treatment = c('control', 'treatment1', 'treatment2', 'control', 'treatment1', 'treatment2','control', 'treatment1', 'treatment2') Sp1 = c(56, 42, 67, 23, 44, 21, 15, 20, 12) Sp2 = c(15, 10, 17, 1, 5, 2, 3, 1,6) Sp3 = c(10, 6, 7, 10, 5, 4, 0, 1, 0) Sp4 = c(9, 6, 4, 8, 13, 5, 2, 1, 0) df = data.frame(Site, Treatment, Sp1, Sp2, Sp3, Sp4)
步骤2:定义分组计算函数
该函数会对每个站点的子数据框执行以下操作:
- 计算站点内所有样本的Bray-Curtis距离矩阵
- 提取每个样本到同站点对照组的距离
- 将对照组自身的距离设为0(可选,因为距离矩阵中自身距离本来就是0)
calculate_bray_to_control <- function(sub_df) { # 计算站点内样本的Bray-Curtis距离矩阵 dist_matrix <- vegdist(sub_df[, 3:6], method = "bray") %>% as.matrix() # 定位对照组的行索引 control_row <- which(sub_df$Treatment == "control") # 提取每个样本到对照组的距离 sub_df$dis.to.control <- dist_matrix[, control_row] # 确保对照组自身距离为0(可选) sub_df$dis.to.control[control_row] <- 0 return(sub_df) }
步骤3:批量处理所有站点
利用group_split按站点拆分数据,再通过map_dfr批量应用计算函数并合并结果:
result_df <- df %>% group_split(Site) %>% map_dfr(calculate_bray_to_control)
输出结果示例
运行代码后,result_df会新增dis.to.control列,格式如下(数值为实际计算值):
> print(result_df) Site Treatment Sp1 Sp2 Sp3 Sp4 dis.to.control 1 A control 56 15 10 9 0.0000000 2 A treatment1 42 10 6 6 0.1206897 3 A treatment2 67 17 7 4 0.1458333 4 B control 23 1 10 8 0.0000000 5 B treatment1 44 5 5 13 0.2352941 6 B treatment2 21 2 4 5 0.1627907 7 C control 15 3 0 2 0.0000000 8 C treatment1 20 1 1 1 0.1702128 9 C treatment2 12 6 0 0 0.2000000
该方法无需手动拆分站点,可高效处理15+处理组、20+站点的大规模数据。
内容的提问来源于stack exchange,提问作者Felipe Albornoz
相关产品推荐
相关产品推荐

