You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python生成随机(x,y)点以测试K-means聚类算法?

Hey there! Looks like you're in the middle of generating random points for testing K-means clustering—nice start! Let me help you wrap up that code and even optimize it to create more meaningful clusters that'll work better for your K-means tests.

First, let's fix & complete your original code

If you want to stick with your initial approach of defining each cluster separately, here's the finished version:

import numpy as np

N = 100
# Generate x-coordinates for 3 clusters
random_x0 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4))
random_x1 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4))
random_x2 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4))
# Generate corresponding y-coordinates (finished your cut-off line!)
random_y0 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4))
random_y1 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4))
random_y2 = np.random.randn(N) + (np.random.randint(0, 100) * np.random.randint(1, 4))

# Combine all points into a single dataset ready for K-means
dataset = np.concatenate([
    np.column_stack((random_x0, random_y0)),
    np.column_stack((random_x1, random_y1)),
    np.column_stack((random_x2, random_y2))
])

But let's make it better for K-means testing

Your original code works, but it has a small issue: each cluster's center is totally random every time you run it, which can lead to clusters overlapping so much that K-means can't distinguish them. Here's a more robust approach where we define fixed (or controlled-random) cluster centers first, then generate noisy points around them—this creates clear, distinct clusters perfect for testing:

import numpy as np
import matplotlib.pyplot as plt  # Optional, for visualizing your clusters

# Configuration
num_points_per_cluster = 100
num_clusters = 3

# Generate fixed cluster centers (you can tweak the ranges to adjust spread)
cluster_centers = np.random.randint(0, 100, size=(num_clusters, 2)) * np.random.randint(1, 4, size=(num_clusters, 2))

# Generate points for each cluster
clusters = []
for center in cluster_centers:
    # Add Gaussian noise to center coordinates to create cluster spread
    x_points = np.random.randn(num_points_per_cluster) + center[0]
    y_points = np.random.randn(num_points_per_cluster) + center[1]
    clusters.append(np.column_stack((x_points, y_points)))

# Combine all clusters into one dataset
dataset = np.vstack(clusters)

# Optional: Visualize the clusters to verify
plt.scatter(dataset[:, 0], dataset[:, 1], alpha=0.7)
plt.scatter(cluster_centers[:, 0], cluster_centers[:, 1], c='red', marker='x', s=150, label='Cluster Centers')
plt.legend()
plt.show()

This approach has two big wins:

  • You get consistent, distinct clusters every time (or controlled randomness if you want it)
  • It's scalable—just change num_clusters to generate more clusters without writing repetitive code

内容的提问来源于stack exchange,提问作者Birdman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:58:38