如何用Python生成随机(x,y)点以测试K-means聚类算法?
Hey there! Looks like you're in the middle of generating random points for testing K-means clustering—nice start! Let me help you wrap up that code and even optimize it to create more meaningful clusters that'll work better for your K-means tests.
First, let's fix & complete your original code
If you want to stick with your initial approach of defining each cluster separately, here's the finished version:
import numpy as np N = 100 # Generate x-coordinates for 3 clusters random_x0 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4)) random_x1 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4)) random_x2 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4)) # Generate corresponding y-coordinates (finished your cut-off line!) random_y0 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4)) random_y1 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4)) random_y2 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4)) # Combine all points into a single dataset ready for K-means dataset = np.concatenate([ np.column_stack((random_x0, random_y0)), np.column_stack((random_x1, random_y1)), np.column_stack((random_x2, random_y2)) ])
But let's make it better for K-means testing
Your original code works, but it has a small issue: each cluster's center is totally random every time you run it, which can lead to clusters overlapping so much that K-means can't distinguish them. Here's a more robust approach where we define fixed (or controlled-random) cluster centers first, then generate noisy points around them—this creates clear, distinct clusters perfect for testing:
import numpy as np import matplotlib.pyplot as plt # Optional, for visualizing your clusters # Configuration num_points_per_cluster = 100 num_clusters = 3 # Generate fixed cluster centers (you can tweak the ranges to adjust spread) cluster_centers = np.random.randint(0, 100, size=(num_clusters, 2)) * np.random.randint(1, 4, size=(num_clusters, 2)) # Generate points for each cluster clusters = [] for center in cluster_centers: # Add Gaussian noise to center coordinates to create cluster spread x_points = np.random.randn(num_points_per_cluster) + center[0] y_points = np.random.randn(num_points_per_cluster) + center[1] clusters.append(np.column_stack((x_points, y_points))) # Combine all clusters into one dataset dataset = np.vstack(clusters) # Optional: Visualize the clusters to verify plt.scatter(dataset[:, 0], dataset[:, 1], alpha=0.7) plt.scatter(cluster_centers[:, 0], cluster_centers[:, 1], c='red', marker='x', s=150, label='Cluster Centers') plt.legend() plt.show()
This approach has two big wins:
- You get consistent, distinct clusters every time (or controlled randomness if you want it)
- It's scalable—just change
num_clustersto generate more clusters without writing repetitive code
内容的提问来源于stack exchange,提问作者Birdman

