优化同一sf对象中多要素的凸包生成效率
Great question! When dealing with thousands of SF features, for loops can get pretty slow—especially with repeated bind_rows() calls that copy your entire data frame every iteration. Let's refactor your workflow using dplyr grouping + purrr::pmap to make it vectorized and way more efficient. Here's how to adapt your existing logic:
Step 1: Load Data & Prep Grouped Structure
First, we'll load your data as before, then organize it so each UID's original polygon and associated points are grouped together (this avoids repeated filtering in loops):
library(sf) library(dplyr) library(purrr) library(concaveman) # Load and prepare data (same as your original code) download.file("https://drive.google.com/uc?export=download&id=1-I4F2NYvFWkNqy7ASFNxnyrwr_wT0lGF" , destfile="ProximityAreas.zip") unzip("ProximityAreas.zip") Proximity_Areas <- st_read("Proximity_Areas.gpkg") Proximity_Areas_Points <- Proximity_Areas %>% st_cast("POINT") # Group data by UID: each row holds one UID's original polygon and its points uid_groups <- Proximity_Areas %>% left_join( Proximity_Areas_Points %>% group_by(UID) %>% nest(points = geometry), # Nest points into a list column per UID by = "UID" ) %>% select(UID, geometry, points)
Step 2: Encapsulate Processing Logic in a Function
We'll turn your loop's step-by-step logic into a reusable function. This makes the code cleaner and easier to debug:
process_uid_group <- function(uid, original_poly, points) { # Generate convex hull (concavity=1 = strict convex hull) convex_hull <- st_make_valid(concaveman(points, concavity = 1, length_threshold = 0)) # Clip convex hull to original polygon boundary clipped_convex <- st_intersection(convex_hull, original_poly) %>% mutate(UID = uid) %>% select(UID) # Handle geometry collections (extract and union polygons) if (st_geometry_type(clipped_convex) == "GEOMETRYCOLLECTION") { clipped_convex <- clipped_convex %>% st_collection_extract("POLYGON") %>% group_by(UID) %>% summarize(geometry = st_union(geometry)) %>% ungroup() } # Convert non-polygon geometries (points/lines) to polygons via buffer if (!(st_geometry_type(clipped_convex) %in% c("POLYGON", "MULTIPOLYGON"))) { end_cap_style <- ifelse(st_geometry_type(clipped_convex) == "POINT", "SQUARE", "FLAT") clipped_convex <- clipped_convex %>% st_buffer(dist = 12.5, endCapStyle = end_cap_style) %>% st_cast("MULTIPOLYGON") %>% st_intersection(original_poly) %>% select(UID) } # Ensure geometry is MULTIPOLYGON for consistency if (st_geometry_type(clipped_convex) == "POLYGON") { clipped_convex <- st_cast(clipped_convex, "MULTIPOLYGON") } # Fix any invalid geometries if (!st_is_valid(clipped_convex)) { clipped_convex <- st_make_valid(clipped_convex) } return(clipped_convex) }
Step 3: Process All Groups with purrr::pmap
Instead of looping through each UID manually, we'll use pmap to apply our function to every group in parallel (conceptually), then combine the results:
# Process all UIDs and combine results into a single SF object Proximity_Convex <- uid_groups %>% mutate(convex_result = pmap(list(UID, geometry, points), process_uid_group)) %>% select(convex_result) %>% unnest(convex_result) %>% st_as_sf()
Why This Is Faster Than Your Original Loop
- No repeated data copying: Your original loop uses
bind_rows()every iteration, which creates a full copy of your growing data frame each time. For thousands of features, this is a massive performance drain. - Optimized vectorized operations:
dplyr::nest()andpurrr::pmap()handle iteration in a way that minimizes R's for-loop overhead, making it much faster for large datasets. - Cleaner maintainability: Encapsulating logic in a function makes it easier to tweak edge cases (like buffer sizes or geometry fixes) without rewriting the entire loop.
内容的提问来源于stack exchange,提问作者Chris

