基于Pandas数据集的列间K-Means聚类分析需求咨询
Got it, let's walk through how to get your K-Means clustering done on this hourly time series dataset—where we're grouping columns (each column is a time series) into clusters. Here's a straightforward, step-by-step solution that'll output exactly the format you asked for:
First, we need to make sure our data is loaded correctly, then transpose it because K-Means treats rows as samples—and in your case, each column (A, B, C, etc.) is the "sample" we want to cluster, with each hourly timestamp being a feature.
import pandas as pd from sklearn.cluster import KMeans from sklearn.preprocessing import StandardScaler # Load your dataset (swap in your actual file path or data source) df = pd.read_csv('your_time_series_data.csv', index_col=0, parse_dates=True) # Transpose so each row becomes a column's full time series df_transposed = df.T
Time series columns often have different value scales, which can throw off K-Means (since it uses distance metrics). Standardizing the data ensures every feature carries equal weight:
# Scale features to have mean 0 and standard deviation 1 scaler = StandardScaler() scaled_time_series = scaler.fit_transform(df_transposed)
If you don't already know how many clusters you want, use the elbow method to find the optimal K:
import matplotlib.pyplot as plt inertia_values = [] k_candidates = range(1, 10) for k in k_candidates: kmeans = KMeans(n_clusters=k, random_state=42) kmeans.fit(scaled_time_series) inertia_values.append(kmeans.inertia_) # Plot the elbow curve to spot where inertia stops dropping sharply plt.plot(k_candidates, inertia_values, 'bx-') plt.xlabel('Number of Clusters (K)') plt.ylabel('Inertia (Within-Cluster Sum of Squares)') plt.title('Elbow Method to Find Optimal K') plt.show()
Look for the "elbow" point on the plot—this is your best K (e.g., if inertia flattens out at K=3, go with 3 clusters).
Now let's run the clustering and attach labels to each column:
# Replace with your chosen K (from elbow method or your requirement) n_clusters = 3 kmeans = KMeans(n_clusters=n_clusters, random_state=42) cluster_labels = kmeans.fit_predict(scaled_time_series) # Add cluster labels back to our transposed dataframe df_transposed['Cluster'] = cluster_labels
Finally, we'll group the columns by their cluster and print them in the format you want:
# Loop through each cluster and print the assigned columns for cluster_id in range(n_clusters): # Get all columns in this cluster cluster_cols = df_transposed[df_transposed['Cluster'] == cluster_id].index.tolist() # Format and print print(f"Cluster{cluster_id + 1}: Columns {', '.join(cluster_cols)}")
Quick Notes to Keep in Mind
- Handle Missing Values: K-Means can't work with NaNs—use
df.fillna(method='ffill')or linear interpolation if you have gaps in your time series. - Alternative Scaling: If your data is all positive, you might use
MinMaxScalerinstead ofStandardScaler—but standardization is safer for most cases. - Validate Clusters: Use the silhouette score (
sklearn.metrics.silhouette_score) to check how well-separated your clusters are.
内容的提问来源于stack exchange,提问作者Uhalden

