如何在R中按最近距离合并数据集?百万级数据优化求助
Great question—your brute-force double loop works fine for small datasets, but with 1.5 million rows in A and 5,000 in B, you’re stuck doing 7.5 billion distance calculations. That’s why it’s running so slow. Let’s replace that with efficient, scalable approaches built for this exact kind of spatial matching:
Option 1: Use RANN for KD-Tree Based Nearest Neighbor Search
The RANN package uses a KD-Tree to drastically reduce the number of distance calculations needed. This is the fastest approach for large planar point datasets, with a time complexity of O(M log N) instead of O(M*N).
First install and load the package:
install.packages("RANN") library(RANN)
Then process your data:
# Keep your original data loading code A <- read.table(header = TRUE, text = " Name x y city 50.3 4.2 farm 14.8 8.6 lake 18.7 9.8 mountain 44 9.8") B <- read.table(header = TRUE, text = " Temp x y 18 50.7 6.2 17.3 20 11 15 15 9") # Extract coordinate matrices (required for RANN) A_coords <- as.matrix(A[, c("x", "y")]) B_coords <- as.matrix(B[, c("x", "y")]) # Find the 1 nearest neighbor in B for every point in A nn_results <- nn2(data = B_coords, query = A_coords, k = 1) # Build your final dataset C C <- data.frame( Name = A$Name, Temp = B$Temp[nn_results$nn.idx[, 1]], # Match the nearest B's Temp Distance = nn_results$nn.dists[, 1] # Pre-calculated distance ) print(C)
This will run in a fraction of the time of your loop—KD-Trees are designed exactly for this kind of large-scale nearest neighbor problem.
Option 2: Use sf for Spatial Joins (Great for Geographic Data)
If your coordinates are real geographic coordinates (latitude/longitude), the sf package is the industry standard. It uses spatial indexing to optimize searches, and handles spherical distance calculations correctly if you set the right coordinate reference system (CRS).
install.packages("sf") library(sf) # Convert data to spatial objects (adjust CRS to match your data; 4326 = WGS84 lat/lng) A_sf <- st_as_sf(A, coords = c("x", "y"), crs = 4326) B_sf <- st_as_sf(B, coords = c("x", "y"), crs = 4326) # Perform a nearest-neighbor spatial join joined <- st_join(A_sf, B_sf, join = st_nearest_feature) # Calculate distances and build your result C <- data.frame( Name = joined$Name, Temp = joined$Temp, Distance = st_distance(A_sf, joined$geometry, by_element = TRUE) %>% as.numeric() ) print(C)
This method is more flexible for geographic data, and still way faster than your original loop.
Why Your Original Code Is So Slow
Your double loop uses a brute-force approach: every point in A checks every point in B to find the closest one. With 1.5M * 5K = 7.5B calculations, even fast CPUs will take hours to finish. KD-Trees and spatial indexes avoid this by organizing the data to skip unnecessary distance checks.
内容的提问来源于stack exchange,提问作者Frennetix

