基于聚类在多边形内生成新GPS点的技术咨询
聚类算法是否适用于多边形内生成新GPS点的需求?
我有一个由GPS坐标构成的多边形(比如波兰国界),以及该多边形内30个分布不均的已存在GPS点。我的目标是在这个多边形内生成10个新的GPS点,但尝试K-Means、DBSCAN等算法多日都没成功:要么只生成1个点,要么修改了原有点位,要么生成的点不在多边形内,而且完全没有均匀分布的迹象。我懂聚类的原理以及它对数据集的作用,但不知道怎么把它用到这个场景里,想请教:聚类是否适合这类需求?我是不是对聚类的应用有误解?
我尝试的代码
import numpy as np import matplotlib.pyplot as plt from sklearn.cluster import KMeans from shapely.geometry import Point, Polygon # Define the coordinates for the polygon coords = [(51.20750669474411, 14.771064752102495), (50.3723916491531, 11.814788193533916), (48.50440302296322, 13.967327062741662), (48.534998749507885, 16.951318714046824), (49.58224848458308, 19.057665762026932), (50.43716140927117, 17.7365796749166), (51.166973917905295, 15.399273520798316), (51.20750669474411, 14.771064752102495)] polygon = Polygon(coords) # Define the existing points within the polygon existing_points = [(50.17959204532467, 15.933794312781023), (49.73773753981636, 13.625912823974723), (49.21735566602075, 16.84855409350198), (50.62311076818633, 14.563920611496346), (50.11205389160936, 12.752308482828356), (49.628634641977655, 18.19787451760077), (49.67842445515853, 16.801569465294627), (50.03339498666959, 14.48862122510739), (49.447870172525306, 14.285816210548106), (50.115288565936275, 12.567911717003131), (50.25183321926949, 15.835340641617428), (49.79213754282857, 15.495520622255166)] # Extract the coordinates of the existing points existing_coords = np.array([(p[0], p[1]) for p in existing_points]) # Create a scatter plot of the existing points fig, ax = plt.subplots(figsize=(8, 6)) ax.scatter(existing_coords[:, 0], existing_coords[:, 1], c='blue') # Generate new points using k-means clustering n_clusters = 9 kmeans = KMeans(n_clusters=n_clusters) kmeans.fit(existing_coords) new_coords = kmeans.cluster_centers_ # Check if the new points are within the polygon new_points = [] for coord in new_coords: point = Point(coord[0], coord[1]) if polygon.contains(point): new_points.append(coord) # Add the new points to the scatter plot ax.scatter(np.array(new_points)[:, 0], np.array(new_points)[:, 1], c='red') # Show the plot plt.show()
运行结果示例
- 当
n_clusters = 2时:生成的红色聚类中心点集中在少数区域,完全没有均匀分布的效果 - 当
n_clusters = 9时:多个红色中心点重叠或分布零散,部分点超出多边形范围
结论:聚类算法并非这类需求的最优选择,你确实存在应用场景的误解
为什么现有聚类方法行不通?
- K-Means的本质是「数据驱动的聚类中心」:它的聚类中心是现有数据点的均值,完全依赖输入数据的分布。如果原有点分布不均,比如大部分点集中在某一块,K-Means的中心也会扎堆在那里,不可能自动填补多边形内的空白区域,更做不到均匀分布。
- DBSCAN的局限性:它是基于密度的聚类,只会在数据密集区域生成聚类,空白区域根本不会产生新点,反而会把稀疏点归为噪声,完全不符合你要生成新点的需求。
- 边界问题:聚类算法不会主动考虑多边形边界约束,计算出的中心很容易落在多边形外,事后过滤的方式只能被动丢弃,无法保证生成足够数量的有效点。
适合你需求的方法
你要的是在多边形约束内生成均匀/合理分布的新点,这类需求属于「空间采样」或「点集补全」,推荐以下思路:
- 规则网格采样 + 多边形过滤:在多边形外接矩形内生成均匀网格点,再过滤掉多边形外的点,最后挑选你需要的数量(或结合原有点的分布,优先填补空白区域)。
- 泊松圆盘采样:能生成均匀且避免过近的点,天然适合空间分布需求,同时可以通过多边形约束过滤无效点。
- 基于 Voronoi 图的补全:用原有点生成Voronoi图,找出多边形内没有被覆盖的Voronoi区域,在这些区域的中心生成新点,实现空白区域的填补。
修正后的思路示例(规则网格采样)
import numpy as np import matplotlib.pyplot as plt from shapely.geometry import Point, Polygon # 定义多边形和原有点 coords = [(51.20750669474411, 14.771064752102495), (50.3723916491531, 11.814788193533916), (48.50440302296322, 13.967327062741662), (48.534998749507885, 16.951318714046824), (49.58224848458308, 19.057665762026932), (50.43716140927117, 17.7365796749166), (51.166973917905295, 15.399273520798316), (51.20750669474411, 14.771064752102495)] polygon = Polygon(coords) existing_points = [(50.17959204532467, 15.933794312781023), (49.73773753981636, 13.625912823974723), (49.21735566602075, 16.84855409350198), (50.62311076818633, 14.563920611496346), (50.11205389160936, 12.752308482828356), (49.628634641977655, 18.19787451760077), (49.67842445515853, 16.801569465294627), (50.03339498666959, 14.48862122510739), (49.447870172525306, 14.285816210548106), (50.115288565936275, 12.567911717003131), (50.25183321926949, 15.835340641617428), (49.79213754282857, 15.495520622255166)] existing_coords = np.array(existing_points) # 生成规则网格点 min_lon, min_lat, max_lon, max_lat = polygon.bounds # 调整网格密度,确保能生成足够多的候选点 lon_steps = np.linspace(min_lon, max_lon, 20) lat_steps = np.linspace(min_lat, max_lat, 20) grid_points = [] for lon in lon_steps: for lat in lat_steps: p = Point(lon, lat) if polygon.contains(p): grid_points.append((lon, lat)) grid_points = np.array(grid_points) # 筛选空白区域的点:计算与原有点的最小距离,优先选距离远的 def calc_min_distance(gp, existing_points): return np.min(np.sqrt((gp[0]-existing_points[:,0])**2 + (gp[1]-existing_points[:,1])**2)) distances = [calc_min_distance(gp, existing_coords) for gp in grid_points] # 按距离从大到小排序,取前10个 sorted_indices = np.argsort(distances)[::-1] new_points = grid_points[sorted_indices[:10]] # 可视化 fig, ax = plt.subplots(figsize=(8, 6)) ax.scatter(existing_coords[:, 0], existing_coords[:, 1], c='blue', label='原有点') ax.scatter(new_points[:, 0], new_points[:, 1], c='red', label='新生成点') # 绘制多边形边界 polygon_coords = np.array(coords) ax.plot(polygon_coords[:, 0], polygon_coords[:, 1], c='green', label='多边形边界') ax.legend() plt.show()
内容的提问来源于stack exchange,提问作者Martin Kavka
相关产品推荐
相关产品推荐

