You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大数据集下基于距离的权重矩阵代码优化技术求助

Optimizing nb2blocknb for large datasets (24077 locations, 569k individuals)

I'm trying to create a spatial weight matrix based on distance. My code works fine on small datasets, but fails to execute when processing a large dataset with 24077 locations and 569424 individuals—specifically, the nb2blocknb function is causing the issue. How can I optimize my code for large-scale data?

Original code:

# load all survey data
DHS <- read.csv("Daten/final.csv")
attach(DHS)
# define coordinates matrix
coormat <- cbind(DHS$location, DHS$lon_s, DHS$lat_s)
coorm <- cbind(DHS$lon_s, DHS$lat_s)
colnames(coormat) <- c("location", "lon_s", "lat_s")
coo <- cbind(unique(coormat))
c <- as.data.frame(coo)
coor <- cbind(c$lon_s, c$lat_s)
# get a list with neighbored locations that are within 50 km distance
neighbor <- dnearneigh(coor, d1 = 0, d2 = 50, row.names=c$location, longlat=TRUE, bound=c("GE", "LE"))
# get neighborhood list on individual level
nb <- nb2blocknb(neighbor, as.character(DHS$location))
# weight matrix in list format
nbweights.lw <- nb2listw(nb, style="B", zero.policy=TRUE)

Solutions to Optimize Your Code

Great question—handling large spatial datasets efficiently often boils down to cutting unnecessary memory overhead and replacing slow loop-based functions with vectorized operations. Here’s how you can get your code working smoothly with the large dataset:

1. Simplify Data Preprocessing to Reduce Memory Bloat

Your original code creates redundant data objects (like coormat, coorm, coo) that waste critical memory. Let’s streamline this step to only keep what’s needed:

library(spdep)
library(data.table) # For fast, memory-efficient data manipulation

# Load data with data.table (faster and uses less memory than read.csv)
DHS <- fread("Daten/final.csv")

# Extract unique locations and coordinates directly (no redundant copies)
unique_locations <- unique(DHS[, .(location, lon_s, lat_s)])
coor <- as.matrix(unique_locations[, .(lon_s, lat_s)])
row.names(coor) <- unique_locations$location

2. Replace nb2blocknb with a Vectorized Alternative

nb2blocknb relies on slow loop-based mapping for large individual-level datasets. Instead, use data.table to map each individual to their location’s neighbors in a vectorized way—this is orders of magnitude faster:

# Compute neighborhood list for unique locations (same as before)
neighbor <- dnearneigh(coor, d1 = 0, d2 = 50, longlat = TRUE, bound = c("GE", "LE"))

# Convert neighbor list to a lookup table for fast joins
neighbor_dt <- data.table(
  location = row.names(coor),
  neighbors = I(lapply(neighbor, function(x) row.names(coor)[x]))
)

# Map each individual to their location's neighbors (vectorized operation)
DHS <- DHS[neighbor_dt, on = "location", neighbors := i.neighbors]

# Convert the neighbors column to a valid `nb` object (required for nb2listw)
nb <- lapply(DHS$neighbors, function(x) which(unique_locations$location %in% x))
class(nb) <- "nb"
attr(nb, "region.id") <- DHS$location
attr(nb, "call") <- match.call()

3. Critical Memory-Saving Habits

  • Ditch attach(): It duplicates data in memory and causes unexpected bugs. Use explicit column references (e.g., DHS$location) instead.
  • Clean up unused objects: After creating coor and unique_locations, run rm(coormat, coorm, coo, c) followed by gc() to free up memory immediately.
  • Stick to data.table: It uses less memory than base R data.frame and has faster join/filter operations.

4. Final Weight Matrix Creation

With the optimized nb object, create your weight matrix as before—it should now handle the large dataset without issues:

nbweights.lw <- nb2listw(nb, style = "B", zero.policy = TRUE)

Why This Works

  • Vectorized operations in data.table avoid the slow loop logic that cripples nb2blocknb for large datasets.
  • Cutting redundant data objects reduces memory usage drastically, preventing crashes or slowdowns.
  • Explicit lookup mapping eliminates the overhead of nb2blocknb’s internal data conversions.

内容的提问来源于stack exchange,提问作者Kerstin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:54:33