基于Python的Open OrderIDs与Filled OrderIDs分层聚类匹配咨询
Hey there, nice job getting your data loaded and merged already—let's walk through exactly how to implement hierarchical clustering and match each open order to its most similar filled order subsets. I'll break this into actionable steps with code examples you can adapt to your data.
1. First: Data Prep & Feature Engineering
Hierarchical clustering relies on numerical, scaled features, so we need to clean and standardize your data first. Here's how to do it:
Step 1: Split Open/Filled Orders
First, separate your merged dataset into open and filled orders (I'll assume you have a flag column like is_open to distinguish them):
import pandas as pd # Split your merged data open_orders = merged_df[merged_df['is_open'] == True] filled_orders = merged_df[merged_df['is_open'] == False]
Step 2: Preprocess Features
Handle categorical variables (like product category or region) and scale numerical features (like quantity or price) so they don't dominate the clustering:
from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.compose import ColumnTransformer # Replace these with your actual feature columns numeric_features = ['Quantity', 'UnitPrice', 'EstimatedDeliveryDays'] categorical_features = ['ProductCategory', 'CustomerRegion'] # Build a preprocessing pipeline preprocessor = ColumnTransformer( transformers=[ ('scale_numeric', StandardScaler(), numeric_features), ('encode_categorical', OneHotEncoder(drop='first'), categorical_features) ]) # Fit the pipeline to all orders (to ensure consistent feature space) processed_all = preprocessor.fit_transform(merged_df.drop(['OrderID', 'is_open'], axis=1)) # Convert processed features back to a DataFrame for easier handling processed_df = pd.DataFrame( processed_all.toarray(), columns=preprocessor.get_feature_names_out(), index=merged_df['OrderID'] ) # Split processed features into open/filled subsets open_processed = processed_df.loc[open_orders['OrderID']] filled_processed = processed_df.loc[filled_orders['OrderID']]
2. Implement Hierarchical Clustering
You have two solid approaches here: cluster all orders together and match open orders to filled orders in the same cluster, or calculate pairwise distances to pull the most similar filled orders for each open order.
Approach 1: Cluster All Orders & Match by Cluster
This groups similar orders into clusters, so every open order's cluster will contain its most similar filled orders:
from sklearn.cluster import AgglomerativeClustering # Initialize the clustering model # Use ward linkage to minimize within-cluster variance; adjust distance_threshold to control cluster size cluster_model = AgglomerativeClustering( n_clusters=None, distance_threshold=0.6, # Lower = smaller, more similar clusters linkage='ward', metric='euclidean' ) # Assign cluster labels to all orders merged_df['cluster_label'] = cluster_model.fit_predict(processed_all) # Match each open order to filled orders in the same cluster order_cluster_matches = {} for _, open_order in open_orders.iterrows(): cluster_id = open_order['cluster_label'] similar_filled = filled_orders[filled_orders['cluster_label'] == cluster_id]['OrderID'].tolist() order_cluster_matches[open_order['OrderID']] = similar_filled # Check results print("Cluster-based matches:") for open_id, matches in order_cluster_matches.items(): print(f"Open Order {open_id}: {matches}")
If you want to visualize clusters to tune the distance_threshold, use a dendrogram:
from scipy.cluster.hierarchy import dendrogram, linkage import matplotlib.pyplot as plt # Calculate linkage matrix linked = linkage(processed_all, 'ward') # Plot dendrogram plt.figure(figsize=(12, 6)) dendrogram(linked, orientation='top', distance_sort='descending', show_leaf_counts=True) plt.title('Hierarchical Clustering Dendrogram') plt.xlabel('Order IDs') plt.ylabel('Euclidean Distance') plt.show()
Cut the dendrogram at a height that gives you logically sized clusters, then set that height as your distance_threshold.
Approach 2: Pairwise Distance Matching
If you want to pull the top N most similar filled orders for each open order (instead of entire clusters), calculate pairwise distances:
from scipy.spatial.distance import cdist # Calculate Euclidean distance between each open order and all filled orders distance_matrix = cdist(open_processed, filled_processed, metric='euclidean') # Convert to DataFrame for readability distance_df = pd.DataFrame( distance_matrix, index=open_processed.index, columns=filled_processed.index ) # Get top 5 most similar filled orders (adjust N as needed) top_n = 5 order_top_matches = {} for open_id in distance_df.index: # Sort distances and take the smallest N closest_filled = distance_df.loc[open_id].sort_values().head(top_n).index.tolist() order_top_matches[open_id] = closest_filled print(f"Top {top_n} matches per open order:") for open_id, matches in order_top_matches.items(): print(f"Open Order {open_id}: {matches}")
3. Tuning & Validation Tips
- Adjust Distance Metrics: If your features are high-dimensional or mixed-type, try
cosinesimilarity instead of Euclidean distance, orcityblock(Manhattan) for robust results. - Evaluate Cluster Quality: Use the Silhouette Score to test different cluster counts—scores closer to 1 mean better-defined clusters:
from sklearn.metrics import silhouette_score for n_clusters in range(2, 10): model = AgglomerativeClustering(n_clusters=n_clusters, linkage='ward') labels = model.fit_predict(processed_all) score = silhouette_score(processed_all, labels) print(f"n_clusters={n_clusters}, Silhouette Score={score:.4f}") - Feature Selection: If you have too many features, use PCA to reduce dimensionality or filter features by their correlation to order similarity—this reduces noise and speeds up clustering.
内容的提问来源于stack exchange,提问作者Bryan Hammond

