如何基于经纬度匹配合并R语言中ex与clim数据集?
Hey there! Let's figure out how to match each location in your ex dataset to the closest geographic point in the clim dataset, so we can merge the climate data correctly. Here are two reliable approaches using R:
First, let's load your sample data
First, make sure you have the datasets ready (copy-paste this to replicate your setup):
ex <- data.frame(lat = c(55, 60, 40), long = c(6, 6, 10)) clim <- structure(list(lat = c(55.047, 55.097, 55.146, 55.004, 55.054, 55.103, 55.153, 55.202, 55.252, 55.301), long = c(6.029, 6.0171, 6.0051, 6.1269, 6.1151, 6.1032, 6.0913, 6.0794, 6.0675, 6.0555 ), alt = c(0.033335, 0.033335, 0.033335, 0.033335, 0.033335, 0.033335, 0.033335, 0.033335, 0.033335, 0.033335), x = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), y = c(1914, 1907.3, 1901.8, 1921.1, 1914.1, 1908.3, 1902.4, 1896, 1889.8, 1884)), row.names = c(NA, 10L), class = "data.frame")
Approach 1: Using Base R + geosphere (accurate spherical distance)
Since we're working with geographic coordinates (latitude/longitude), using spherical distance is more accurate than Euclidean distance (especially for large latitudinal spans). The geosphere package has a handy distHaversine function that calculates this distance in meters.
# Install the package if you haven't already # install.packages("geosphere") library(geosphere) # Iterate over each row in ex to find the closest clim point matched_data <- lapply(1:nrow(ex), function(row_idx) { # Get the current location from ex current_point <- ex[row_idx, c("long", "lat")] # Calculate distance from this point to all points in clim distances <- distHaversine(current_point, clim[, c("long", "lat")]) # Find the index of the closest point in clim closest_clim_idx <- which.min(distances) # Merge the ex row with the corresponding clim data (exclude duplicate lat/long) cbind(ex[row_idx, ], clim[closest_clim_idx, setdiff(names(clim), c("lat", "long"))]) }) # Convert the list of rows into a single data frame matched_data <- do.call(rbind, matched_data) # View the result print(matched_data)
Approach 2: Tidyverse + purrr (modern R workflow)
If you prefer a pipe-based workflow, this uses dplyr and purrr to achieve the same result in a more readable way:
# Install the tidyverse if needed # install.packages("tidyverse") library(tidyverse) library(geosphere) matched_data_tidy <- ex %>% rowwise() %>% mutate( # Find the index of the closest clim point closest_idx = which.min(distHaversine(c(long, lat), clim[, c("long", "lat")])), # Pull in the climate variables from the closest clim row alt = clim$alt[closest_idx], x = clim$x[closest_idx], y = clim$y[closest_idx] ) %>% ungroup() %>% select(-closest_idx) # Remove the helper column # View the tidy result print(matched_data_tidy)
Result
Both approaches will give you the exact output you're expecting:
lat long alt x y 1 55 6 0.033335 0 1914.0 2 60 6 0.033335 0 1884.0 3 40 10 0.033335 0 1921.1
A quick note: For the third row (lat=40, long=10), the closest point in clim is row 4 (y=1921.1) because it has the smallest spherical distance to (40,10) compared to all other clim points.
内容的提问来源于stack exchange,提问作者Mateusz1981

