You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python的Open OrderIDs与Filled OrderIDs分层聚类匹配咨询

Hey there, nice job getting your data loaded and merged already—let's walk through exactly how to implement hierarchical clustering and match each open order to its most similar filled order subsets. I'll break this into actionable steps with code examples you can adapt to your data.

1. First: Data Prep & Feature Engineering

Hierarchical clustering relies on numerical, scaled features, so we need to clean and standardize your data first. Here's how to do it:

Step 1: Split Open/Filled Orders

First, separate your merged dataset into open and filled orders (I'll assume you have a flag column like is_open to distinguish them):

import pandas as pd

# Split your merged data
open_orders = merged_df[merged_df['is_open'] == True]
filled_orders = merged_df[merged_df['is_open'] == False]

Step 2: Preprocess Features

Handle categorical variables (like product category or region) and scale numerical features (like quantity or price) so they don't dominate the clustering:

from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer

# Replace these with your actual feature columns
numeric_features = ['Quantity', 'UnitPrice', 'EstimatedDeliveryDays']
categorical_features = ['ProductCategory', 'CustomerRegion']

# Build a preprocessing pipeline
preprocessor = ColumnTransformer(
    transformers=[
        ('scale_numeric', StandardScaler(), numeric_features),
        ('encode_categorical', OneHotEncoder(drop='first'), categorical_features)
    ])

# Fit the pipeline to all orders (to ensure consistent feature space)
processed_all = preprocessor.fit_transform(merged_df.drop(['OrderID', 'is_open'], axis=1))

# Convert processed features back to a DataFrame for easier handling
processed_df = pd.DataFrame(
    processed_all.toarray(),
    columns=preprocessor.get_feature_names_out(),
    index=merged_df['OrderID']
)

# Split processed features into open/filled subsets
open_processed = processed_df.loc[open_orders['OrderID']]
filled_processed = processed_df.loc[filled_orders['OrderID']]

2. Implement Hierarchical Clustering

You have two solid approaches here: cluster all orders together and match open orders to filled orders in the same cluster, or calculate pairwise distances to pull the most similar filled orders for each open order.

Approach 1: Cluster All Orders & Match by Cluster

This groups similar orders into clusters, so every open order's cluster will contain its most similar filled orders:

from sklearn.cluster import AgglomerativeClustering

# Initialize the clustering model
# Use ward linkage to minimize within-cluster variance; adjust distance_threshold to control cluster size
cluster_model = AgglomerativeClustering(
    n_clusters=None,
    distance_threshold=0.6,  # Lower = smaller, more similar clusters
    linkage='ward',
    metric='euclidean'
)

# Assign cluster labels to all orders
merged_df['cluster_label'] = cluster_model.fit_predict(processed_all)

# Match each open order to filled orders in the same cluster
order_cluster_matches = {}
for _, open_order in open_orders.iterrows():
    cluster_id = open_order['cluster_label']
    similar_filled = filled_orders[filled_orders['cluster_label'] == cluster_id]['OrderID'].tolist()
    order_cluster_matches[open_order['OrderID']] = similar_filled

# Check results
print("Cluster-based matches:")
for open_id, matches in order_cluster_matches.items():
    print(f"Open Order {open_id}: {matches}")

If you want to visualize clusters to tune the distance_threshold, use a dendrogram:

from scipy.cluster.hierarchy import dendrogram, linkage
import matplotlib.pyplot as plt

# Calculate linkage matrix
linked = linkage(processed_all, 'ward')

# Plot dendrogram
plt.figure(figsize=(12, 6))
dendrogram(linked,
           orientation='top',
           distance_sort='descending',
           show_leaf_counts=True)
plt.title('Hierarchical Clustering Dendrogram')
plt.xlabel('Order IDs')
plt.ylabel('Euclidean Distance')
plt.show()

Cut the dendrogram at a height that gives you logically sized clusters, then set that height as your distance_threshold.

Approach 2: Pairwise Distance Matching

If you want to pull the top N most similar filled orders for each open order (instead of entire clusters), calculate pairwise distances:

from scipy.spatial.distance import cdist

# Calculate Euclidean distance between each open order and all filled orders
distance_matrix = cdist(open_processed, filled_processed, metric='euclidean')

# Convert to DataFrame for readability
distance_df = pd.DataFrame(
    distance_matrix,
    index=open_processed.index,
    columns=filled_processed.index
)

# Get top 5 most similar filled orders (adjust N as needed)
top_n = 5
order_top_matches = {}
for open_id in distance_df.index:
    # Sort distances and take the smallest N
    closest_filled = distance_df.loc[open_id].sort_values().head(top_n).index.tolist()
    order_top_matches[open_id] = closest_filled

print(f"Top {top_n} matches per open order:")
for open_id, matches in order_top_matches.items():
    print(f"Open Order {open_id}: {matches}")

3. Tuning & Validation Tips

  • Adjust Distance Metrics: If your features are high-dimensional or mixed-type, try cosine similarity instead of Euclidean distance, or cityblock (Manhattan) for robust results.
  • Evaluate Cluster Quality: Use the Silhouette Score to test different cluster counts—scores closer to 1 mean better-defined clusters:
    from sklearn.metrics import silhouette_score
    
    for n_clusters in range(2, 10):
        model = AgglomerativeClustering(n_clusters=n_clusters, linkage='ward')
        labels = model.fit_predict(processed_all)
        score = silhouette_score(processed_all, labels)
        print(f"n_clusters={n_clusters}, Silhouette Score={score:.4f}")
    
  • Feature Selection: If you have too many features, use PCA to reduce dimensionality or filter features by their correlation to order similarity—this reduces noise and speeds up clustering.

内容的提问来源于stack exchange,提问作者Bryan Hammond

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:07:42