如何将物种存在数据经纬度匹配至WorldClim 2.5弧分分辨率以适配MaxEnt建模
匹配物种存在记录与WorldClim 2.5弧分网格的实操步骤
要把不同精度的物种存在记录匹配到WorldClim 2.5弧分(对应5位小数经纬度,约5km²)的网格单元,进而生成存在/缺失数据集用于MaxEnt建模,按以下步骤操作:
1. 明确WorldClim网格的坐标规则
2.5弧分等于0.0416667度(2.5/60),每个网格单元的中心经纬度保留5位小数。每个网格的边界范围为:
- 经度边界:
中心经度 ± 0.0208333(0.0416667/2) - 纬度边界:
中心纬度 ± 0.0208333
所有匹配操作都要基于这个边界范围,而非直接四舍五入坐标(避免因精度差异导致的网格匹配错误)。
2. 匹配物种记录到对应网格单元
方法一:用R语言(生态建模常用工具)
# 加载所需包 library(raster) library(dplyr) # 加载任意一个WorldClim 2.5弧分的bioclim图层(用于获取网格信息) bioclim_raster <- raster("bio1.tif") # 替换为你的bioclim文件路径 # 读取物种存在记录(假设数据框为sp_occ,包含lon、lat列) sp_occ <- read.csv("species_occurrences.csv") # 转换为空间点对象,确保坐标系与WorldClim一致(WGS84,EPSG:4326) sp_points <- SpatialPoints(sp_occ[,c("lon","lat")], proj4string = CRS(proj4string(bioclim_raster))) # 过滤掉欧洲范围外的点(根据你的研究区域调整经纬度范围) europe_extent <- extent(-10, 30, 35, 70) # 示例欧洲范围,按需修改 sp_points <- crop(sp_points, europe_extent) # 获取每个点对应的网格单元ID,再转为网格中心经纬度 cell_ids <- cellFromXY(bioclim_raster, sp_points) grid_centers <- xyFromCell(bioclim_raster, cell_ids) # 合并到物种数据框,去重(同一个网格只保留一条存在记录) sp_occ_matched <- cbind(sp_occ, grid_centers) %>% distinct(x, y, .keep_all = TRUE) %>% rename(grid_lon = x, grid_lat = y) %>% mutate(presence = 1)
方法二:用Python语言
import rasterio import geopandas as gpd import pandas as pd from shapely.geometry import Point # 加载WorldClim栅格 with rasterio.open("bio1.tif") as src: bio_transform = src.transform bio_crs = src.crs # 读取物种存在记录 sp_occ = pd.read_csv("species_occurrences.csv") # 转为GeoDataFrame,确保坐标系为WGS84 sp_occ = gpd.GeoDataFrame( sp_occ, geometry=gpd.points_from_xy(sp_occ.lon, sp_occ.lat), crs="EPSG:4326" ) # 确保坐标系一致,若不一致则转换 if sp_occ.crs != bio_crs: sp_occ = sp_occ.to_crs(bio_crs) # 过滤欧洲范围外的点(示例范围,按需调整) europe_bounds = [-10, 35, 30, 70] # min_lon, min_lat, max_lon, max_lat sp_occ = sp_occ.cx[europe_bounds[0]:europe_bounds[2], europe_bounds[1]:europe_bounds[3]] # 计算每个点对应的栅格行列号,再转为中心经纬度 def point_to_grid_center(point, transform): row, col = ~transform * (point.x, point.y) row, col = int(round(row)), int(round(col)) center_lon, center_lat = transform * (col + 0.5, row + 0.5) return round(center_lon, 5), round(center_lat, 5) sp_occ[["grid_lon", "grid_lat"]] = sp_occ.geometry.apply( lambda p: point_to_grid_center(p, bio_transform) ).apply(pd.Series) # 去重,标记存在 sp_occ_matched = sp_occ.drop_duplicates(subset=["grid_lon", "grid_lat"]).assign(presence=1)
3. 生成完整的存在/缺失数据集
- 提取欧洲范围内所有WorldClim网格的中心经纬度:
- 在R中可以用
rasterToPoints(bioclim_raster)获取所有网格的中心坐标,再过滤到欧洲范围; - 在Python中可以基于栅格的行列号,结合转换参数计算所有网格的中心坐标。
- 在R中可以用
- 将匹配后的物种存在网格与全量网格做左连接:
- 存在记录的网格标记为
1,无存在记录的网格标记为0,最终得到每个网格单元的存在/缺失标签。
- 存在记录的网格标记为
关键注意事项
- 必须确保物种记录与WorldClim栅格的坐标系完全一致(默认都是WGS84,EPSG:4326),否则会出现匹配错误;
- 不要直接对物种记录的经纬度做四舍五入,必须基于网格边界或行列号计算对应中心坐标,避免精度偏差导致的网格错配;
- 过滤掉研究区域(欧洲)外的物种记录,减少无效计算。
内容的提问来源于stack exchange,提问作者Insect_biologist
相关产品推荐
相关产品推荐

