基于经纬度数据的R语言K-means聚类实现(采用Haversine距离)
Hey there! Let's work through your problem—since you're dealing with a small dataset (only 20 points) of latitude/longitude coordinates, it makes total sense that density-based clustering (like DBSCAN) didn't pan out. Density methods rely on dense regions of points, which can be tough with such a small sample. K-means (or its more robust cousin, PAM) paired with Haversine distance is a great alternative here, and I'll walk you through exactly how to implement it.
Why Density Clustering Failed for Your Data
Density-based algorithms like DBSCAN are super sensitive to parameters (eps and minPts) and need enough points in dense clusters to work well. With only 20 data points, you likely don't have enough density to form meaningful groups, hence the poor results. K-means-style methods are way better suited for smaller datasets where you can define the number of clusters upfront.
Step 1: Prepare Your Data & Calculate Haversine Distance
First, we'll use the geosphere package to compute the Haversine distance matrix—this accounts for the spherical nature of Earth, unlike Euclidean distance which would be totally inaccurate for lat/long data.
Assuming your dataset looks like this (swap this with your actual data):
# Sample dataset (replace with your lat/long data) lat_long_data <- data.frame( longitude = c(-74.0060, -118.2437, -0.1278, 139.6917, 100.5018), latitude = c(40.7128, 34.0522, 51.5074, 35.6895, 13.7563) )
Now compute the Haversine distance matrix:
# Load required package library(geosphere) # Convert data to matrix (geosphere expects lon/lat order) coords_matrix <- as.matrix(lat_long_data[, c("longitude", "latitude")]) # Calculate Haversine distance matrix (units: meters) haversine_dist_matrix <- distm(coords_matrix, fun = distHaversine)
Step 2: Run K-means-style Clustering with Custom Distance
Standard kmeans() in R doesn't support custom distance matrices, so we'll use PAM (Partitioning Around Medoids) from the cluster package. PAM works similarly to K-means but uses actual data points as cluster centers (medoids) instead of virtual centroids, which is way more intuitive for geographic data.
library(cluster) # Choose number of clusters (k) – start with a guess, then refine k <- 3 # Adjust this based on your data's structure # Run PAM clustering using the Haversine distance matrix pam_cluster <- pam(haversine_dist_matrix, k = k, diss = TRUE) # Add cluster labels to your original dataset lat_long_data$cluster <- as.factor(pam_cluster$clustering)
How to Choose the Right k?
Use the elbow method to find the optimal number of clusters:
# Calculate total within-cluster distance for k=1 to k=5 within_cluster_dist <- sapply(1:5, function(k_val) { pam(haversine_dist_matrix, k = k_val, diss = TRUE)$objective }) # Plot the elbow curve plot(1:5, within_cluster_dist, type = "b", xlab = "Number of Clusters (k)", ylab = "Total Within-Cluster Haversine Distance", main = "Elbow Method for Optimal k")
Look for the "elbow" point where the distance stops decreasing sharply—that's your ideal k.
Step 3: Visualize Clusters on Google Maps
To plot your clusters on Google Maps, we'll use the ggmap package (note: you'll need a Google Maps API key to use this). Alternatively, leaflet offers a free interactive map option if you don't want to set up an API key.
Option 1: ggmap (Google Maps Background)
library(ggmap) # Register your Google Maps API key first (get one from Google Cloud Console) # register_google(key = "YOUR_GOOGLE_API_KEY") # Get a map centered on your data map_center <- c(lon = mean(lat_long_data$longitude), lat = mean(lat_long_data$latitude)) google_map <- get_googlemap(center = map_center, zoom = 4) # Plot clusters on the map ggmap(google_map) + geom_point(data = lat_long_data, aes(x = longitude, y = latitude, color = cluster), size = 4, alpha = 0.8) + labs(color = "Cluster") + theme_minimal()
Option 2: Leaflet (Free Interactive Map)
If you don't want to use a Google API key, Leaflet is a great alternative:
library(leaflet) # Create interactive map with clusters leaflet(lat_long_data) %>% addTiles() # Uses OpenStreetMap by default (you can add Google Maps tiles too) %>% addCircleMarkers(lng = ~longitude, lat = ~latitude, color = ~cluster, radius = 6, stroke = FALSE, fillOpacity = 0.8) %>% addLegend("bottomright", pal = colorFactor("viridis", lat_long_data$cluster), values = ~cluster, title = "Cluster")
Key Takeaways
- For small lat/long datasets, density clustering often underperforms—K-means/PAM is a better fit
- Always use Haversine (or Vincenty) distance for geographic data, never Euclidean
- PAM is easier to implement with custom distance matrices than modifying K-means directly
- Use the elbow method to pick the right number of clusters
- Visualize with ggmap (Google Maps) or Leaflet for clear, actionable results
内容的提问来源于stack exchange,提问作者Nikita Agarwal

