You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于经纬度数据的R语言K-means聚类实现(采用Haversine距离)

K-means Clustering with Haversine Distance for Lat/Long Data (R)

Hey there! Let's work through your problem—since you're dealing with a small dataset (only 20 points) of latitude/longitude coordinates, it makes total sense that density-based clustering (like DBSCAN) didn't pan out. Density methods rely on dense regions of points, which can be tough with such a small sample. K-means (or its more robust cousin, PAM) paired with Haversine distance is a great alternative here, and I'll walk you through exactly how to implement it.

Why Density Clustering Failed for Your Data

Density-based algorithms like DBSCAN are super sensitive to parameters (eps and minPts) and need enough points in dense clusters to work well. With only 20 data points, you likely don't have enough density to form meaningful groups, hence the poor results. K-means-style methods are way better suited for smaller datasets where you can define the number of clusters upfront.

Step 1: Prepare Your Data & Calculate Haversine Distance

First, we'll use the geosphere package to compute the Haversine distance matrix—this accounts for the spherical nature of Earth, unlike Euclidean distance which would be totally inaccurate for lat/long data.

Assuming your dataset looks like this (swap this with your actual data):

# Sample dataset (replace with your lat/long data)
lat_long_data <- data.frame(
  longitude = c(-74.0060, -118.2437, -0.1278, 139.6917, 100.5018),
  latitude = c(40.7128, 34.0522, 51.5074, 35.6895, 13.7563)
)

Now compute the Haversine distance matrix:

# Load required package
library(geosphere)

# Convert data to matrix (geosphere expects lon/lat order)
coords_matrix <- as.matrix(lat_long_data[, c("longitude", "latitude")])

# Calculate Haversine distance matrix (units: meters)
haversine_dist_matrix <- distm(coords_matrix, fun = distHaversine)

Step 2: Run K-means-style Clustering with Custom Distance

Standard kmeans() in R doesn't support custom distance matrices, so we'll use PAM (Partitioning Around Medoids) from the cluster package. PAM works similarly to K-means but uses actual data points as cluster centers (medoids) instead of virtual centroids, which is way more intuitive for geographic data.

library(cluster)

# Choose number of clusters (k) – start with a guess, then refine
k <- 3  # Adjust this based on your data's structure

# Run PAM clustering using the Haversine distance matrix
pam_cluster <- pam(haversine_dist_matrix, k = k, diss = TRUE)

# Add cluster labels to your original dataset
lat_long_data$cluster <- as.factor(pam_cluster$clustering)

How to Choose the Right k?

Use the elbow method to find the optimal number of clusters:

# Calculate total within-cluster distance for k=1 to k=5
within_cluster_dist <- sapply(1:5, function(k_val) {
  pam(haversine_dist_matrix, k = k_val, diss = TRUE)$objective
})

# Plot the elbow curve
plot(1:5, within_cluster_dist, type = "b", 
     xlab = "Number of Clusters (k)", 
     ylab = "Total Within-Cluster Haversine Distance",
     main = "Elbow Method for Optimal k")

Look for the "elbow" point where the distance stops decreasing sharply—that's your ideal k.

Step 3: Visualize Clusters on Google Maps

To plot your clusters on Google Maps, we'll use the ggmap package (note: you'll need a Google Maps API key to use this). Alternatively, leaflet offers a free interactive map option if you don't want to set up an API key.

Option 1: ggmap (Google Maps Background)

library(ggmap)

# Register your Google Maps API key first (get one from Google Cloud Console)
# register_google(key = "YOUR_GOOGLE_API_KEY")

# Get a map centered on your data
map_center <- c(lon = mean(lat_long_data$longitude), lat = mean(lat_long_data$latitude))
google_map <- get_googlemap(center = map_center, zoom = 4)

# Plot clusters on the map
ggmap(google_map) +
  geom_point(data = lat_long_data, 
             aes(x = longitude, y = latitude, color = cluster), 
             size = 4, alpha = 0.8) +
  labs(color = "Cluster") +
  theme_minimal()

Option 2: Leaflet (Free Interactive Map)

If you don't want to use a Google API key, Leaflet is a great alternative:

library(leaflet)

# Create interactive map with clusters
leaflet(lat_long_data) %>%
  addTiles()  # Uses OpenStreetMap by default (you can add Google Maps tiles too) %>%
  addCircleMarkers(lng = ~longitude, lat = ~latitude, 
                   color = ~cluster, radius = 6, stroke = FALSE, fillOpacity = 0.8) %>%
  addLegend("bottomright", pal = colorFactor("viridis", lat_long_data$cluster), 
            values = ~cluster, title = "Cluster")

Key Takeaways

  • For small lat/long datasets, density clustering often underperforms—K-means/PAM is a better fit
  • Always use Haversine (or Vincenty) distance for geographic data, never Euclidean
  • PAM is easier to implement with custom distance matrices than modifying K-means directly
  • Use the elbow method to pick the right number of clusters
  • Visualize with ggmap (Google Maps) or Leaflet for clear, actionable results

内容的提问来源于stack exchange,提问作者Nikita Agarwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:09:37