如何将物种共现分析的SES_C.score结果整合到R数据集对应行
问题:将物种共现SES_Cscore值匹配到数据集(无顺序物种对)
需求说明
需要将基于data2计算的物种共现SES_Cscore值整合到data1中,核心难点是无顺序物种对的匹配——比如RC与BABO的组合,无论谁在Species列、谁在Association列,都要对应同一个SES_Cscore值,普通merge函数无法实现这种匹配。
现有数据集
data1(待添加Cscore列)
Species Association Day month year season cover evi <fctr> <chr> <dbl> <dbl> <dbl> <chr> <chr> <chr> 21 BW RC 22 1 2 dry Evergreen edge 22 MANG RC 22 1 2 dry Evergreen edge 23 RC BW 22 1 2 dry Evergreen edge 24 RC SKS 22 1 2 dry Evergreen edge 25 RC NA 22 1 2 dry Mixed edge 26 SKS MANG 22 1 2 dry Evergreen edge 27 RC NA 30 1 2 dry Mixed edge 28 SKS NA 30 1 2 dry Mixed edge 29 RC NA 30 1 2 dry Mixed edge 30 RC BABO 30 1 2 dry Mixed edge
data2(用于计算SES_Cscore)
group_id RC BW SKS BABO MANG <chr> <dbl><dbl><dbl><dbl><dbl> 2-1-15-Deciduous.dry_470 1 0 0 0 0 2-1-15-Evergreen.dry_1850 0 1 0 0 0 2-1-15-Evergreen.dry_2020 1 1 0 0 0 2-1-23-Semi-Deciduous.dry_1000 0 0 1 0 0 2-1-23-Semi-Deciduous.dry_1310 0 0 1 0 0 2-1-23-Open.dry_1745 1 0 0 0 0 2-1-25-Evergreen.dry_1805 1 1 0 0 0 2-1-25-Evergreen.dry_2050 1 0 0 0 0 2-1-29-Mixed.dry_750 1 0 0 1 0 2-1-29-Evergreen.dry_1958 1 0 0 0 1
计算SES_Cscore的代码:
data2 <- data2[,2:6] nperm <- 1000 outpath <- getwd() Cscore <- ecospat.Cscore(data2, nperm, outpath, verbose = T)
计算得到的SES_Cscore结果表
Sp1 Sp2 obs.C.score exp.C.score SES_Cscore pval_less pval_greater <chr> <chr> <int> <int> <dbl> <dbl> <dbl> RC BW 5 5 -0.8079579 0.4845155 0.9520480 RC SKS 14 6 1.5873256 1.0000000 0.2317682 RC BABO 0 7 -1.0829308 0.4605395 1.0000000 RC MANG 0 0 -1.1117101 0.4475524 1.0000000 BW SKS 6 6 0.5844538 1.0000000 0.7442557 BW BABO 3 3 0.3143282 1.0000000 0.9100899 BW MANG 3 3 0.3914705 1.0000000 0.8671329 SKS BABO 2 2 0.2699785 1.0000000 0.9320679 SKS MANG 2 2 0.2804816 1.0000000 0.9270729 BABO MANG 1 1 0.1903501 1.0000000 0.9650350
期望输出格式
Species Association Day month year season cover evi Cscore <fctr> <chr> <dbl> <dbl> <dbl> <chr> <chr> <chr> <dbl> 21 BW RC 22 1 2 dry Evergreen edge -0.8079579 22 MANG RC 22 1 2 dry Evergreen edge -1.1117101 23 RC BW 22 1 2 dry Evergreen edge -0.8079579 24 RC SKS 22 1 2 dry Evergreen edge 1.5873256 25 RC NA 22 1 2 dry Mixed edge NA 26 SKS MANG 22 1 2 dry Evergreen edge 0.2804816 27 RC NA 30 1 2 dry Mixed edge NA 28 SKS NA 30 1 2 dry Mixed edge NA 29 RC NA 30 1 2 dry Mixed edge NA 30 RC BABO 30 1 2 dry Mixed edge -1.0829308
解决方案
核心思路是统一物种对的顺序:无论原数据中两个物种的顺序如何,都将它们按固定规则排序(比如字母顺序)生成匹配键,再进行关联。
步骤1:处理Cscore结果表,生成无顺序匹配键
library(dplyr) # 对Cscore结果中的Sp1和Sp2按字母排序,生成统一的匹配键 Cscore_matched <- Cscore %>% mutate( # 将物种对转为固定顺序的字符串,比如"RC_BW"和"BW_RC"都转为"BW_RC" pair_key = pmap_chr(list(Sp1, Sp2), ~paste(sort(c(..1, ..2)), collapse = "_")) ) %>% # 每个匹配键只保留一条SES_Cscore记录 select(pair_key, SES_Cscore) %>% distinct(pair_key, .keep_all = TRUE)
步骤2:处理data1,关联Cscore值
# 为data1生成匹配键,并关联对应的SES_Cscore data1_with_cscore <- data1 %>% mutate( # 仅当Association不为NA时生成匹配键,否则为NA pair_key = case_when( !is.na(Association) ~ pmap_chr(list(Species, Association), ~paste(sort(c(as.character(..1), ..2)), collapse = "_")), TRUE ~ NA_character_ ) ) %>% # 关联Cscore值 left_join(Cscore_matched, by = "pair_key") %>% # 重命名列并清理冗余字段 rename(Cscore = SES_Cscore) %>% select(-pair_key)
运行上述代码后,data1_with_cscore即可得到期望的输出,无顺序物种对会正确匹配到同一个SES_Cscore值,Association为NA的行对应NA的Cscore。
示例数据代码
data1 <- data.frame( Species = factor( c("BW", "MANG", "RC", "RC", "RC", "SKS", "RC", "SKS", "RC", "RC"), levels = c("BABO", "BW", "MANG", "RC", "SKS") ), Association = c("RC", "RC", "BW", "SKS", NA, "MANG", NA, NA, NA, "BABO"), Day = c(22, 22, 22, 22, 22, 22, 30, 30, 30, 30), month = c(1, 1, 1, 1, 1, 1, 1, 1, 1, 1), year = c(2, 2, 2, 2, 2, 2, 2, 2, 2, 2), season = c( "dry", "dry", "dry", "dry", "dry", "dry", "dry", "dry", "dry", "dry" ), cover = c( "Evergreen", "Evergreen", "Evergreen", "Evergreen", "Mixed", "Evergreen", "Mixed", "Mixed", "Mixed", "Mixed" ), evi = c( "edge", "edge", "edge", "edge", "edge", "edge", "edge", "edge", "edge", "edge" ), row.names = 21:30 ) data2 <- tibble::tribble( ~group_id, ~RC, ~BW, ~SKS, ~BABO, ~MANG, "2-1-15-Deciduous.dry_470", 1, 0, 0, 0, 0, "2-1-15-Evergreen.dry_1850", 0, 1, 0, 0, 0, "2-1-15-Evergreen.dry_2020", 1, 1, 0, 0, 0, "2-1-23-Semi-Deciduous.dry_1000", 0, 0, 1, 0, 0, "2-1-23-Semi-Deciduous.dry_1310", 0, 0, 1, 0, 0, "2-1-23-Open.dry_1745", 1, 0, 0, 0, 0, "2-1-25-Evergreen.dry_1805", 1, 1, 0, 0, 0, "2-1-25-Evergreen.dry_2050", 1, 0, 0, 0, 0, "2-1-29-Mixed.dry_750", 1, 0, 0, 1, 0, "2-1-29-Evergreen.dry_1958", 1, 0, 0, 0, 1, )
内容的提问来源于stack exchange,提问作者Marnee Roundtree
相关产品推荐
相关产品推荐

