You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言现有函数并行化:提升凸包计算代码效率的方法

Optimizing Convex Hull Calculation with Parallel Processing in R

First, let's fix and optimize the sequential code to avoid redundant operations, then move to parallel implementations.

Sequential Optimizations

The original code subsets the data twice (once for chull, once for creating sf objects). We can combine these steps and fix the convex hull to be a polygon instead of points:

library(sf)

set.seed(123)
n <- 100000
df <- data.frame(longitude = runif(n, -180, 180),
                 latitude = runif(n, -90, 90),
                 color = sample(c("red", "blue", "green", "orange", "purple", "yellow", "pink", "black", "white", "grey"), n, replace = TRUE))

# Split data by color once to avoid repeated subsetting
color_groups <- split(df, df$color)

# Compute convex hulls and convert to polygon sf objects in one pass
hull_sfs <- lapply(color_groups, function(group) {
  # Get convex hull indices
  hull_indices <- chull(group[, c("longitude", "latitude")])
  # Extract hull points and close the polygon (repeat first point)
  hull_points <- group[hull_indices, ]
  hull_points <- rbind(hull_points, hull_points[1, ])
  # Convert to valid polygon sf object
  st_as_sf(hull_points, coords = c("longitude", "latitude"), crs = 4326) %>%
    st_combine() %>%
    st_cast("POLYGON") %>%
    st_sf(color = unique(group$color))
})

# Combine results and write to shapefile
hull_sf_combined <- do.call(rbind, hull_sfs)
st_write(hull_sf_combined, "hulls_sequential.shp")

Parallel Implementation with parallel::mclapply (macOS/Linux)

mclapply uses fork-based parallelism, which is efficient for Unix-like systems:

library(sf)
library(parallel)

set.seed(123)
n <- 100000
df <- data.frame(longitude = runif(n, -180, 180),
                 latitude = runif(n, -90, 90),
                 color = sample(c("red", "blue", "green", "orange", "purple", "yellow", "pink", "black", "white", "grey"), n, replace = TRUE))

color_groups <- split(df, df$color)
num_cores <- detectCores() - 1 # Leave one core for system processes

# Parallel computation
hull_sfs <- mclapply(color_groups, function(group) {
  hull_indices <- chull(group[, c("longitude", "latitude")])
  hull_points <- group[hull_indices, ]
  hull_points <- rbind(hull_points, hull_points[1, ])
  st_as_sf(hull_points, coords = c("longitude", "latitude"), crs = 4326) %>%
    st_combine() %>%
    st_cast("POLYGON") %>%
    st_sf(color = unique(group$color))
}, mc.cores = num_cores)

hull_sf_combined <- do.call(rbind, hull_sfs)
st_write(hull_sf_combined, "hulls_parallel_mclapply.shp")

Parallel Implementation with foreach and doSNOW (Cross-Platform)

This method works on Windows, macOS, and Linux using socket clusters:

library(sf)
library(foreach)
library(doSNOW)
library(parallel)

set.seed(123)
n <- 100000
df <- data.frame(longitude = runif(n, -180, 180),
                 latitude = runif(n, -90, 90),
                 color = sample(c("red", "blue", "green", "orange", "purple", "yellow", "pink", "black", "white", "grey"), n, replace = TRUE))

color_groups <- split(df, df$color)
num_cores <- detectCores() - 1

# Create and register cluster
cl <- makeCluster(num_cores)
registerDoSNOW(cl)

# Optional progress bar
pb <- txtProgressBar(max = length(color_groups), style = 3)
progress <- function(n) setTxtProgressBar(pb, n)
opts <- list(progress = progress)

# Parallel loop with foreach
hull_sfs <- foreach(group = color_groups, 
                    .packages = "sf", # Load sf on worker nodes
                    .options.snow = opts) %dopar% {
  hull_indices <- chull(group[, c("longitude", "latitude")])
  hull_points <- group[hull_indices, ]
  hull_points <- rbind(hull_points, hull_points[1, ])
  st_as_sf(hull_points, coords = c("longitude", "latitude"), crs = 4326) %>%
    st_combine() %>%
    st_cast("POLYGON") %>%
    st_sf(color = unique(group$color))
}

# Clean up
close(pb)
stopCluster(cl)

hull_sf_combined <- do.call(rbind, hull_sfs)
st_write(hull_sf_combined, "hulls_parallel_foreach.shp")

Key Notes

  • Polygon Correction: The original code outputs convex hull vertices as points; we convert them to valid closed polygons, which is the standard representation of convex hulls.
  • Redundant Subsetting: Splitting the data once with split() avoids repeated subsetting, which improves efficiency even before parallelization.
  • Reproducibility: Setting set.seed() ensures consistent results. For socket clusters (doSNOW), if your loop uses random numbers, you'll need to set worker-specific seeds (not needed here).
  • Core Count: Using detectCores() -1 prevents overloading your system by leaving one core free for background tasks.

内容的提问来源于stack exchange,提问作者stats_noob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 07:10:23