如何优化R语言中空间稀疏化的运行速度?
问题背景
我正在开展一个需构建约30000个物种分布模型(SDM)的项目,目前的瓶颈是提升空间稀疏化(针对存在大量分布点的物种)的速度,此前一直使用SpThin包完成该操作。我认为SpThin以保留最多分布点为设计目标,但对于分布点极多的物种而言这并非必需,因此尝试了现有工具并探索新思路。
加载依赖包
library(terra) library(spThin) library(enmSdmX) library(microbenchmark)
示例数据测试
先使用简单数据集测试不同方法的速度:
example <- structure(list(x = c(1.5, 2.5, 2, 5.5, 7.5), y = c(1.5, 2.5, 2, 5.5, 7.5)), class = "data.frame", row.names = c(NA, -5L)) example$ID <- 1:nrow(example) example$Sp <- "A" example_spat <- vect(example, crs = "+proj=longlat +datum=WGS84", geom = c("x", "y"))
创建的分布点及80公里缓冲区内的重叠情况见对应图示,使用microbenchmark测试性能:
Test <- microbenchmark::microbenchmark( A = enmSdmX::geoThin(example_spat, minDist = 80000), B = enmSdmX::geoThin(example_spat, minDist = 80000, random = T), c = spThin::thin(loc.data = as.data.frame(example), thin.par = 80, reps = 1, write.files = F, write.log.file = F, lat.col = "x", long.col = "y", spec.col = "Sp"), times = 100)
测试结果显示enmSdmX的随机模式速度最快,但该结果在大数据集下会反转。
大数据集测试
使用自研包SDMWorkflows获取测试数据:
remotes::install_github("Sustainscapes/SDMWorkflows") library(SDMWorkflows) Presences <- SDMWorkflows::GetOccs(Species = c("Abies concolor"), WriteFile = FALSE, limit = 2000) Cleaned <- clean_presences(Presences[[1]]) spat_vect <- terra::vect(as.data.frame(Cleaned), geom=c("decimalLongitude", "decimalLatitude"), crs = "+proj=longlat +datum=WGS84")
再次进行性能测试:
Test <- microbenchmark::microbenchmark( A = enmSdmX::geoThin(spat_vect, minDist = 10000), B = enmSdmX::geoThin(spat_vect, minDist = 10000, random = T), c = spThin::thin(loc.data = as.data.frame(example), thin.par = 10, reps = 1, write.files = F, write.log.file = F, lat.col = "x", long.col = "y", spec.col = "Sp"), times = 20)
结果显示SpThin速度最快,但这正是我处理大分布点物种时的性能瓶颈,自研2-3个新函数仍未实现提速。
内容的提问来源于stack exchange,提问作者Derek Corcoran
相关产品推荐
相关产品推荐

