You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas数据集的列间K-Means聚类分析需求咨询

Got it, let's walk through how to get your K-Means clustering done on this hourly time series dataset—where we're grouping columns (each column is a time series) into clusters. Here's a straightforward, step-by-step solution that'll output exactly the format you asked for:

Step 1: Load and Reshape Your Data

First, we need to make sure our data is loaded correctly, then transpose it because K-Means treats rows as samples—and in your case, each column (A, B, C, etc.) is the "sample" we want to cluster, with each hourly timestamp being a feature.

import pandas as pd
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler

# Load your dataset (swap in your actual file path or data source)
df = pd.read_csv('your_time_series_data.csv', index_col=0, parse_dates=True)

# Transpose so each row becomes a column's full time series
df_transposed = df.T
Step 2: Preprocess the Data

Time series columns often have different value scales, which can throw off K-Means (since it uses distance metrics). Standardizing the data ensures every feature carries equal weight:

# Scale features to have mean 0 and standard deviation 1
scaler = StandardScaler()
scaled_time_series = scaler.fit_transform(df_transposed)
Step 3: Pick the Right Number of Clusters (Optional but Smart)

If you don't already know how many clusters you want, use the elbow method to find the optimal K:

import matplotlib.pyplot as plt

inertia_values = []
k_candidates = range(1, 10)

for k in k_candidates:
    kmeans = KMeans(n_clusters=k, random_state=42)
    kmeans.fit(scaled_time_series)
    inertia_values.append(kmeans.inertia_)

# Plot the elbow curve to spot where inertia stops dropping sharply
plt.plot(k_candidates, inertia_values, 'bx-')
plt.xlabel('Number of Clusters (K)')
plt.ylabel('Inertia (Within-Cluster Sum of Squares)')
plt.title('Elbow Method to Find Optimal K')
plt.show()

Look for the "elbow" point on the plot—this is your best K (e.g., if inertia flattens out at K=3, go with 3 clusters).

Step 4: Run K-Means and Assign Clusters

Now let's run the clustering and attach labels to each column:

# Replace with your chosen K (from elbow method or your requirement)
n_clusters = 3

kmeans = KMeans(n_clusters=n_clusters, random_state=42)
cluster_labels = kmeans.fit_predict(scaled_time_series)

# Add cluster labels back to our transposed dataframe
df_transposed['Cluster'] = cluster_labels
Step 5: Generate Your Desired Output Format

Finally, we'll group the columns by their cluster and print them in the format you want:

# Loop through each cluster and print the assigned columns
for cluster_id in range(n_clusters):
    # Get all columns in this cluster
    cluster_cols = df_transposed[df_transposed['Cluster'] == cluster_id].index.tolist()
    # Format and print
    print(f"Cluster{cluster_id + 1}: Columns {', '.join(cluster_cols)}")

Quick Notes to Keep in Mind

  • Handle Missing Values: K-Means can't work with NaNs—use df.fillna(method='ffill') or linear interpolation if you have gaps in your time series.
  • Alternative Scaling: If your data is all positive, you might use MinMaxScaler instead of StandardScaler—but standardization is safer for most cases.
  • Validate Clusters: Use the silhouette score (sklearn.metrics.silhouette_score) to check how well-separated your clusters are.

内容的提问来源于stack exchange,提问作者Uhalden

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:24:23