如何通过两个关联DataFrame计算位点性状丰度并生成新数据框?
R语言实现位点-性状丰度计算
依赖包
- 基础R无需额外包;若使用tidyverse风格代码,需安装并加载
dplyr和tidyr
基础R实现代码(高效简洁)
# 加载原始数据 sites_df <- structure(list(BaCo_1 = c(0, 0, 0, 0, 0), BaFa_1 = c(0, 0, 2, 0, 0), BaSl_1 = c(3, 0, 0, 1, 0), BrSl_1 = c(2, 1, 0, 0, 0), CaCo_1 = c(9, 0, 2, 0, 0), CaFa_1 = c(0, 0, 0, 0, 7), CaSl_1 = c(0, 2, 0, 0, 1)), row.names = c("Abla.nota", "Albo.woro", "Aust.coll", "Bibu.Kadj", "cala.sp"), class = "data.frame") traits_df <- structure(list(MFA2 = c(1, 0, 0, 1, 0, 0, 0, 0, 1, 0, 0), MFA4 = c(0, 0, 1, 0, 0, 0, 1, 1, 0, 0, 0), MFA5 = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0), MFA6 = c(0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1), MFA12 = c(0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0), MFA14 = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0), flow1 = c(1, 1, 0, 0, 1, 1, 0, 1, 0, 1, 1), flow2 = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0), flow3 = c(0, 0, 1, 1, 0, 0, 1, 0, 1, 0, 0)), row.names = c("Abla.nota", "Albo.woro", "alel.aust", "Aulo.stri", "Aust.anac", "Aust.coll", "Aust.subt", "bero.sp", "Bibu.Kadj", "Bran.sowe", "cala.sp"), class = "data.frame") # 1. 筛选性状数据,仅保留丰度数据中存在的物种 traits_subset <- traits_df[rownames(sites_df), ] # 2. 矩阵乘法计算:转置丰度矩阵(位点×物种) × 性状矩阵(物种×性状) trait_abundance <- as.data.frame(t(sites_df) %*% traits_subset) # 查看结果 print(trait_abundance)
Tidyverse实现代码(直观易读)
若偏好tidyverse风格的代码,可使用以下方式:
# 首次运行需安装依赖包 # install.packages("tidyverse") library(dplyr) library(tidyr) trait_abundance_tidy <- sites_df %>% # 将行名转为物种列 rownames_to_column("species") %>% # 转换为长格式:物种-位点-丰度 pivot_longer(-species, names_to = "site", values_to = "abundance") %>% # 关联性状数据,仅保留匹配的物种 inner_join(traits_df %>% rownames_to_column("species"), by = "species") %>% # 转换性状列为长格式:物种-位点-性状-是否拥有该性状 pivot_longer(starts_with(c("MFA", "flow")), names_to = "trait", values_to = "has_trait") %>% # 仅保留拥有该性状的记录 filter(has_trait == 1) %>% # 按位点和性状分组,求和丰度 group_by(site, trait) %>% summarise(total_abundance = sum(abundance), .groups = "drop") %>% # 转换回宽格式:位点-性状-总丰度,缺失值填充0 pivot_wider(names_from = trait, values_from = total_abundance, values_fill = 0) %>% # 将位点列转为行名 column_to_rownames("site") # 查看结果 print(trait_abundance_tidy)
结果验证
两种方法生成的结果均与手动示例一致,例如CaCo_1位点的flow1性状值为11,对应Abla.nota的9个个体与Aust.coll的2个个体丰度之和。
内容的提问来源于stack exchange,提问作者Sean
相关产品推荐
相关产品推荐

