You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于距离在Python Pandas中对两类站点数据进行聚类?

How to Cluster Geographic Station Data by Distance in Pandas

Great question! Clustering stations based on geographic distance makes perfect sense for your dataset—here's a step-by-step guide using Pandas and scikit-learn, with practical code examples tailored to your two DataFrames.

Step 1: Combine Your Station DataFrames

First, let's merge your small and main stations into a single DataFrame (we'll use sample main station data since you didn't provide it—replace it with your actual data):

import pandas as pd
import numpy as np

# Your small stations data (from your question)
small_stations = pd.DataFrame({
    'Station_ID': ['dongsi_aq', 'tiantan_aq', 'guanyuan_aq', 'wanshouxigong_aq', 
                   'aotizhongxin_aq', 'nongzhanguan_aq', 'wanliu_aq', 'beibuxinqu_aq', 
                   'zhiwuyuan_aq', 'fengtaihuayuan_aq'],
    'longitude': [116.417, 116.407, 116.339, 116.352, 116.397, 116.461, 116.287, 
                  116.174, 116.207, 116.279],
    'latitude': [39.929, 39.886, 39.929, 39.878, 39.982, 39.937, 39.987, 40.090, 
                 40.002, 39.860]  # Filled missing latitude for fengtaihuayuan_aq
})

# Sample main stations data (replace with your actual main station data)
main_stations = pd.DataFrame({
    'Station_ID': ['main1_aq', 'main2_aq', 'main3_aq', 'main4_aq', 'main5_aq'],
    'longitude': [116.35, 116.42, 116.48, 116.22, 116.55],
    'latitude': [39.91, 39.89, 39.94, 40.05, 39.87]
})

# Combine into one DataFrame for clustering
all_stations = pd.concat([small_stations, main_stations], ignore_index=True)

Step 2: Calculate Pairwise Geographic Distances

Never use Euclidean distance for latitude/longitude—it doesn't account for Earth's curvature. Instead, use the Haversine formula to compute real-world distances in kilometers:

from sklearn.metrics.pairwise import haversine_distances
from math import radians

# Convert coordinates from degrees to radians (required for haversine)
coords_rad = np.radians(all_stations[['latitude', 'longitude']].values)

# Compute distance matrix (results in radians; multiply by Earth's radius ~6371 km to get km)
distance_matrix = haversine_distances(coords_rad) * 6371

Step 3: Cluster Using Distance-Based Algorithms

Choose an algorithm based on your clustering goal:

Option 1: DBSCAN (Density-Based Clustering)

Great for finding arbitrary-shaped clusters (e.g., grouping stations that are within a certain distance of each other):

from sklearn.cluster import DBSCAN

# Tune parameters:
# - eps: Maximum distance (km) between points in the same cluster
# - min_samples: Minimum number of points needed to form a cluster
dbscan = DBSCAN(eps=2, min_samples=2, metric='precomputed')
all_stations['cluster_id'] = dbscan.fit_predict(distance_matrix)

# View results
print(all_stations[['Station_ID', 'cluster_id']])

Option 2: Agglomerative Hierarchical Clustering

Builds clusters by merging the closest pairs of points—use this if you want a fixed number of clusters:

from sklearn.cluster import AgglomerativeClustering

# Cluster into 5 groups (adjust n_clusters to your needs)
agg_clustering = AgglomerativeClustering(
    n_clusters=5, metric='precomputed', linkage='average'
)
all_stations['cluster_id'] = agg_clustering.fit_predict(distance_matrix)

print(all_stations[['Station_ID', 'cluster_id']])

Option 3: Assign Small Stations to Nearest Main Station

If your goal is to map each small station to its closest main station (instead of clustering all together):

from geopy.distance import geodesic

def find_nearest_main(small_row, main_df):
    min_distance = float('inf')
    nearest_main = None
    for _, main_row in main_df.iterrows():
        distance = geodesic(
            (small_row['latitude'], small_row['longitude']),
            (main_row['latitude'], main_row['longitude'])
        ).km
        if distance < min_distance:
            min_distance = distance
            nearest_main = main_row['Station_ID']
    return nearest_main, min_distance

# Apply to small stations
small_stations[['nearest_main', 'distance_km']] = small_stations.apply(
    lambda row: pd.Series(find_nearest_main(row, main_stations)), axis=1
)

print(small_stations[['Station_ID', 'nearest_main', 'distance_km']])

Key Tips

  • Tune parameters: For DBSCAN, adjust eps based on how far apart you want stations to be in the same cluster. For hierarchical clustering, use a dendrogram to pick the optimal number of clusters.
  • Efficiency: For larger datasets, avoid full distance matrices—use BallTree from scikit-learn for faster nearest-neighbor searches.
  • Validation: Visualize clusters on a map (using libraries like folium) to check if the results make geographic sense.

内容的提问来源于stack exchange,提问作者Jiayu Zhang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:51:51