基于SciPy构建时序树状图(Dendrogram)遇技术问题
Got it, let's work through this together. The main issue you're hitting is almost certainly a data orientation mix-up—when you want to cluster entire columns (A-Z) as time series, you need to flip your DataFrame first, because hierarchy.linkage() treats rows as individual samples by default. Your original approach was clustering 8878 time points instead of 26 columns, which is why the dendrogram was unreadable.
Here's a step-by-step solution tailored to your dataset and goal:
Step 1: Fix Data Orientation
First, transpose your DataFrame. This turns each column (A-Z) into a single row (a sample of 8878 time points), which is exactly what we need for clustering columns based on their full-time series behavior.
import pandas as pd import scipy.cluster.hierarchy as sch import matplotlib.pyplot as plt # Assume your original DataFrame is named `df` (rows = timestamps, columns = A-Z) df_transposed = df.T # Now rows = A-Z, columns = timestamps
Step 2: Choose a Time-Series Friendly Distance Metric
Ordinary Euclidean distance works if your time series are perfectly aligned (each timestamp corresponds to the same event). But for most real-world time series, Dynamic Time Warping (DTW) is better—it accounts for shifts or delays in similar patterns.
Option A: DTW (Recommended for Timing Shifts)
Install fastdtw first for efficient calculations, then build a distance matrix for your 26 columns:
from fastdtw import fastdtw from scipy.spatial.distance import euclidean import numpy as np # Build a 26x26 distance matrix for A-Z columns distance_matrix = np.zeros((len(df_transposed), len(df_transposed))) for i in range(len(df_transposed)): for j in range(len(df_transposed)): distance, _ = fastdtw(df_transposed.iloc[i].values, df_transposed.iloc[j].values, dist=euclidean) distance_matrix[i][j] = distance # Create linkage matrix with Ward's method Z = sch.linkage(distance_matrix, method='ward')
Option B: Euclidean Distance (For Strictly Aligned Data)
If your time series are perfectly synchronized, you can skip building a custom distance matrix:
# Directly use transposed DataFrame—linkage auto-calculates Euclidean distance Z = sch.linkage(df_transposed, method='ward')
Step 3: Plot a Readable Dendrogram
Now you can plot a dendrogram that clearly shows how columns A-Z cluster based on their time series patterns:
plt.figure(figsize=(12, 6)) # Plot dendrogram with column labels (A-Z), rotated to avoid overlap dendro = sch.dendrogram( Z, labels=df_transposed.index.values, orientation='top', leaf_rotation=90, leaf_font_size=12 ) plt.title('Hierarchical Clustering of Time Series Columns (A-Z)') plt.xlabel('Columns') plt.ylabel('Distance') plt.tight_layout() plt.show()
Step 4: Advanced: Combine with Heatmap (RSC-Style)
To match the detailed, publication-ready style you mentioned, pair the dendrogram with a time series heatmap. This shows both clustering and raw data patterns:
import seaborn as sns # Get the ordered column names from the dendrogram cluster_order = dendro['ivl'] # Reorder your original DataFrame to match clustering df_clustered = df[cluster_order] # Plot clustermap with column clustering plt.figure(figsize=(15, 8)) g = sns.clustermap( df_clustered, row_cluster=False, # Don't cluster timestamps col_linkage=Z, # Use our precomputed column linkage cmap='viridis', figsize=(15, 8) ) g.fig.suptitle('Time Series Heatmap with Hierarchical Column Clustering', y=1.02) plt.show()
Why Your Original Plot Failed
When you passed the untransposed DataFrame to linkage(), you were trying to cluster 8878 individual time points instead of 26 columns. A dendrogram with 8878 leaf nodes is impossible to read—this fix flips the focus to the column-level time series you actually care about.
内容的提问来源于stack exchange,提问作者A.Solen

