You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过R获取rgbif下载的GBIF occurrence数据的Habitat等属性

如何用R的rgbif包获取GBIF数据中的生境、鉴定备注和分布备注字段?

我已成功用rgbif包下载了GBIF的物种分布数据,但获取到的属性有限,想拿到Habitat(生境)、**Identification remarks(鉴定备注)和Occurrence remarks(分布备注)**字段,请问用R能不能实现?

我使用的代码如下:

library(rgbif)
# 请求数据 
occ_download(pred_and(pred("phylumKey", 35), pred("gadm","ETH")),
             format = "SIMPLE_CSV")

# 用下载编号检查状态,编号来自上一步控制台输出
occ_download_wait("下载编号")

# 加载数据
df <- occ_download_get("下载编号") %>%
  occ_download_import()

查看数据列名后,未找到目标字段:

names(df)
 [1] "gbifID"                           "datasetKey"                       "occurrenceID"                     "kingdom"                          
 [5] "phylum"                           "class"                            "order"                            "family"                           
 [9] "genus"                            "species"                          "infraspecificEpithet"             "taxonRank"                        
[13] "scientificName"                   "verbatimScientificName"           "verbatimScientificNameAuthorship" "countryCode"                      
[17] "locality"                         "stateProvince"                    "occurrenceStatus"                 "individualCount"                  
[21] "publishingOrgKey"                 "decimalLatitude"                  "decimalLongitude"                 "coordinateUncertaintyInMeters"    
[25] "coordinatePrecision"              "elevation"                        "elevationAccuracy"                "depth"                            
[29] "depthAccuracy"                    "eventDate"                        "day"                              "month"                            
[33] "year"                             "taxonKey"                         "speciesKey"                       "basisOfRecord"                    
[37] "institutionCode"                  "collectionCode"                   "catalogNumber"                    "recordNumber"                     
[41] "identifiedBy"                     "dateIdentified"                   "license"                          "rightsHolder"                     
[45] "recordedBy"                       "typeStatus"                       "establishmentMeans"               "lastInterpreted"                  
[49] "mediaType"                        "issue"     

解决方案

可以实现,问题出在你选择的下载格式上:SIMPLE_CSV只包含GBIF定义的核心字段,生境、备注这类扩展字段不在其中。只需要修改格式参数,就能获取完整字段:

方法1:使用完整CSV格式

修改occ_download中的format参数为"CSV"(去掉SIMPLE_前缀),这样下载的文件会包含所有可用字段:

library(rgbif)
# 请求包含所有字段的CSV数据
occ_download(pred_and(pred("phylumKey", 35), pred("gadm","ETH")),
             format = "CSV")

# 等待下载完成
occ_download_wait("你的下载编号")

# 加载完整数据
df_full <- occ_download_get("你的下载编号") %>%
  occ_download_import()

加载后查看列名,就能找到你需要的字段:

  • habitat(对应生境)
  • identificationRemarks(对应鉴定备注)
  • occurrenceRemarks(对应分布备注)

方法2:使用Darwin Core Archive格式(可选)

如果需要更规范的达尔文核心数据集,可以用format = "DWCA",下载的是一个压缩包,解压后其中的occurrence.txt文件包含所有字段。不过这种方式需要手动解压或用代码处理压缩包,比CSV格式稍繁琐。

内容的提问来源于stack exchange,提问作者kl-higgins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 14:20:22