R语言sf包st_union处理大空间数据运行过慢求优化
R语言sf包空间统计优化方案(针对90万条分布点数据)
问题背景
需要实现:将两个sf对象关联后,按shapefile的FYLKESNAVN(县名)与分布点数据的genus(属)分组统计数量:
- 已转EPSG:4326坐标系的shapefile(含县名字段)
- 90万条记录的物种分布点sf对象(含属名字段)
此前尝试情况:
st_join速度正常,但未结合分组统计实现需求st_intersection运行过慢,多次手动终止- 误用
st_union导致运行极慢,已尝试更新R(4.3.3)、相关包、删除.RDATA文件,添加进度条无效
原错误代码:
# Perform the spatial union of shp2 and occ_sf library(sf) library(dplyr) shp2_occsf_union <- st_union(shp2, occ_sf) class(shp2_occsf_union) View(shp2_occsf_union) # Group data grouped_union <- shp2_occsf_union %% group_by(FYLKESNAVN, genus) summarize(occurrences = n()) # View the result View(grouped_union) # Plot the union library("ggplot2") ggplot() + geom_sf(data = shp2) + geom_sf(data = grouped_union, aes(color = occurrences), size = 0.25) + geom_sf(data = shp) + theme_classic()
优化解决方案
1. 优先用st_join实现需求(最高效)
st_join本身速度快,结合分组统计完全能达成目标,无需使用st_union或st_intersection:
library(sf) library(dplyr) # 空间连接:将县属性关联到每个分布点 joined_data <- st_join(occ_sf, shp2, join = st_within) # 按县名+属名分组统计,可选合并同组点几何用于绘图 grouped_stats <- joined_data %>% filter(!is.na(FYLKESNAVN)) %>% # 过滤不在县域内的点(按需选择) group_by(FYLKESNAVN, genus) %>% summarize(occurrences = n(), geometry = st_union(geometry)) %>% # 可选:合并同组点的几何 ungroup() # 绘图 library(ggplot2) ggplot() + geom_sf(data = shp2) + geom_sf(data = grouped_stats, aes(color = occurrences), size = 0.25) + theme_classic()
2. 空间运算前置优化(若需交集类操作)
如果必须使用空间交集,先做以下优化提升速度:
- 转换投影:EPSG:4326是地理坐标系,空间运算效率低,转成当地投影坐标系(如挪威常用的EPSG:25833)运算,完成后再转回4326:
shp2_proj <- st_transform(shp2, 25833) occ_sf_proj <- st_transform(occ_sf, 25833) # 后续运算用投影后的对象,完成后再用st_transform(..., 4326)转回 - 简化多边形几何:若shapefile的多边形过于复杂,用
st_simplify减少运算量:shp2_simple <- st_simplify(shp2, dTolerance = 100) # 100单位(投影后为米),按需调整精度 - 分批次处理:将90万条记录拆分小批次运算后合并:
library(purrr) # 按每1万条拆分批次 batches <- split(occ_sf, ceiling(seq_along(occ_sf$genus)/10000)) # 批量处理每个批次 batch_results <- map(batches, ~st_join(.x, shp2) %>% group_by(FYLKESNAVN, genus) %>% summarize(n = n())) # 合并批次结果并汇总 grouped_stats <- bind_rows(batch_results) %>% group_by(FYLKESNAVN, genus) %>% summarize(occurrences = sum(n)) %>% ungroup()
3. 修正原代码错误
原代码存在两个关键问题:
- 管道符写成了
%%,正确应为%>% - 误用
st_union:该函数默认是合并几何对象,不是关联属性,这是导致运行极慢的核心原因
内容的提问来源于stack exchange,提问作者gn-gia
相关产品推荐
相关产品推荐

