在R中筛选GBIF GRIIS清单入侵物种并获取分布记录
问题概述
需要从GBIF获取多个目标国家的入侵物种分布记录,当前使用ISSG的GRIIS清单时,会下载所有标记为入侵或引入的物种(如印度有2057个),但仅需其中明确标记为Invasive的物种(印度仅282个)。isInvasive字段仅存在于GRIIS数据集的Darwin Core Archive(DwCA)的speciesprofile.txt文件中,手动逐个下载DwCA效率极低,需批量自动化解决方案。
解决方案步骤
1. 批量下载目标国家的GRIIS DwCA文件
通过GBIF API定位目标国家的GRIIS数据集,自动获取DwCA下载链接并批量下载解压:
library(tidyverse) library(rgbif) library(zip) # 设置GBIF凭证 Sys.setenv(GBIF_USER = 'your_username') Sys.setenv(GBIF_PWD = 'your_password') Sys.setenv(GBIF_EMAIL = 'your_email') # 目标国家列表 target_countries <- c('India', 'Malaysia') # 搜索ISSG发布的所有GRIIS checklist数据集 ISSG_datasets <- dataset_search( publishingOrg = 'cdef28b1-db4e-4c58-aa71-3c5238c2d0b5', type = 'checklist' )$data # 筛选目标国家对应的数据集 target_GRIIS <- ISSG_datasets %>% filter(str_detect(datasetTitle, paste(target_countries, collapse = "|"))) # 创建存储DwCA的目录 dir.create("GRIIS_DwCA", showWarnings = FALSE) # 批量下载并解压每个国家的DwCA walk(target_GRIIS$datasetKey, function(key) { # 获取数据集的下载链接 download_info <- dataset_download(key) dwca_url <- download_info$downloadLink # 定义本地存储路径 zip_path <- file.path("GRIIS_DwCA", paste0("griis_", key, ".zip")) extract_dir <- file.path("GRIIS_DwCA", key) # 下载并解压 download.file(dwca_url, zip_path, mode = "wb") unzip(zip_path, exdir = extract_dir) file.remove(zip_path) # 可选:删除压缩包节省空间 })
2. 提取入侵物种并获取GBIF分布记录
从解压后的DwCA中读取speciesprofile.txt,筛选出isInvasive == 'Invasive'的物种,关联GBIF分类学key后批量下载分布记录:
# 安装依赖包(若未安装) install.packages("countrycode") library(countrycode) # 处理单个国家GRIIS数据的函数 process_country_griis <- function(country_dir) { # 读取speciesprofile.txt(注意DwCA默认用制表符分隔) sp_profile <- read.table( file.path(country_dir, "speciesprofile.txt"), header = TRUE, sep = "\t" ) # 筛选入侵物种 invasive_sp <- sp_profile %>% filter(isInvasive == 'Invasive') # 读取taxon文件,关联GBIF nubKey(即taxonKey) taxon_file <- list.files(country_dir, pattern = "^taxon.*\\.txt", full.names = TRUE) taxon_data <- read.table(taxon_file, header = TRUE, sep = "\t") # 关联入侵物种的分类学key invasive_taxa <- invasive_sp %>% left_join(taxon_data, by = "id") %>% filter(!is.na(nubKey)) %>% distinct(nubKey, species = scientificName) # 获取国家ISO2代码(用于GBIF分布筛选) country_name <- str_extract(basename(country_dir), paste(target_countries, collapse = "|")) country_iso2 <- countrycode(country_name, "country.name", "iso2c") # 批量提交GBIF分布下载任务(队列处理避免API限制) if(nrow(invasive_taxa) > 0) { download_tasks <- map(invasive_taxa$nubKey, function(key) { occ_download( pred("taxonKey", key), pred("hasCoordinate", TRUE), pred("country", country_iso2), pred("occurrenceStatus", "PRESENT") ) }) occ_download_queue(download_tasks) } return(invasive_taxa) } # 处理所有下载的DwCA目录 griis_dirs <- list.dirs("GRIIS_DwCA", recursive = FALSE) all_invasive_taxa <- map_dfr(griis_dirs, process_country_griis) # 查看筛选出的入侵物种列表 print(all_invasive_taxa)
关键说明
- DwCA文件格式:GRIIS的DwCA文件默认使用**制表符(\t)**分隔,读取时需指定
sep = "\t",避免解析错误。 - GBIF API限制:使用
occ_download_queue自动处理API速率限制,防止请求被拒绝。 - 分类学关联:通过
id字段关联speciesprofile.txt和taxon.txt,确保获取正确的GBIF物种key。
内容的提问来源于stack exchange,提问作者FAmorim
相关产品推荐
相关产品推荐

