基于给定初始簇的K-Means聚类实现问询:欧氏距离与指定迭代次数
K-Means Clustering with 3 Iterations (Given Initial Centers)
Hey there! Let's work through your K-Means clustering problem step by step. First, I spotted a couple of tiny mistakes in your code:
- The
ycolumn in your DataFrame has a typo: the point(2,2)was incorrectly entered as23instead of2 - You had a partial import (
sklea...) — the correct library name issklearn
Let's start by clarifying the problem details first:
- Raw data points:
(1,1),(1,2),(2,1),(2,2),(3,3),(8,8),(9,8),(8,9),(9,9) - Initial cluster centers:
(1,1)and(2,1) - Constraints: Use Euclidean distance, run exactly 3 iterations
Full Working Code (With Iteration Tracking)
I'll implement this both with a manual iteration approach (for full transparency) and using sklearn (for quick execution), so you can see exactly how centroids update each time:
import pandas as pd import numpy as np from sklearn.cluster import KMeans # Fix the raw data typo first data = {'x': [1,1,2,2,3,8,9,8,9], 'y': [1,2,1,2,3,8,8,9,9]} df = pd.DataFrame(data) # Define initial centers as a numpy array initial_centers = np.array([[1, 1], [2, 1]]) ### Option 1: Manual Iteration (Full Transparency) print("=== Manual 3-Iteration K-Means ===") current_centers = initial_centers.copy() for iter_num in range(3): print(f"\n--- Iteration {iter_num + 1} ---") # Calculate Euclidean distance from each point to both centers distances = np.sqrt(((df - current_centers[:, np.newaxis])**2).sum(axis=2)) # Assign each point to the closest cluster cluster_labels = np.argmin(distances, axis=0) df[f'cluster_iter_{iter_num+1}'] = cluster_labels # Show current cluster assignments print("Cluster Assignments:") print(df[['x', 'y', f'cluster_iter_{iter_num+1}']]) # Update centroids (average of points in each cluster) new_centers = [] for cluster in range(2): cluster_points = df[cluster_labels == cluster][['x', 'y']] new_centroid = cluster_points.mean().values if len(cluster_points) > 0 else current_centers[cluster] new_centers.append(new_centroid) current_centers = np.array(new_centers) print(f"Updated Centroids: {current_centers.round(2)}") ### Option 2: Using sklearn KMeans (Quick Implementation) print("\n=== sklearn K-Means Result ===") kmeans = KMeans( n_clusters=2, init=initial_centers, max_iter=3, n_init=1, # Only use the given initial centers once random_state=42 ) kmeans.fit(df) print(f"Final Centroids: {kmeans.cluster_centers_.round(2)}") print(f"Final Cluster Labels: {kmeans.labels_}")
Step-by-Step Iteration Breakdown
Let's walk through exactly what happens in each iteration:
Iteration 1
- Initial Centroids:
[[1, 1], [2, 1]] - Cluster Assignments:
- Cluster 0:
(1,1),(1,2)(closer to(1,1)) - Cluster 1:
(2,1),(2,2),(3,3),(8,8),(9,8),(8,9),(9,9)(closer to(2,1))
- Cluster 0:
- Updated Centroids:
- Cluster 0:
(1.0, 1.5)(average of its two points) - Cluster 1:
(4.71, 5.71)(average of its seven points)
- Cluster 0:
Iteration 2
- Current Centroids:
[[1.0, 1.5], [4.71, 5.71]] - Cluster Assignments:
- Cluster 0:
(1,1),(1,2),(2,1),(2,2),(3,3)(now closer to the updated Cluster 0 centroid) - Cluster 1:
(8,8),(9,8),(8,9),(9,9)(still closer to Cluster 1 centroid)
- Cluster 0:
- Updated Centroids:
- Cluster 0:
(1.8, 1.8)(average of its five points) - Cluster 1:
(8.5, 8.5)(average of its four points)
- Cluster 0:
Iteration 3
- Current Centroids:
[[1.8, 1.8], [8.5, 8.5]] - Cluster Assignments: No changes from Iteration 2 — all points stay in their current clusters
- Updated Centroids: Identical to Iteration 2 (since cluster membership didn't change, the average remains the same)
Final Clustering Result
After 3 iterations, the clusters stabilize to:
- Cluster 0:
(1,1),(1,2),(2,1),(2,2),(3,3) - Cluster 1:
(8,8),(9,8),(8,9),(9,9)
内容的提问来源于stack exchange,提问作者Shivam
相关产品推荐
相关产品推荐

