You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于高程对二维网格数据进行连续聚类(最小簇规模4)

二维网格高程连续聚类(最小簇大小≥4)

核心思路:区域生长算法

区域生长是适配这类连续区域+属性相似性聚类需求的最优方案之一,核心逻辑:

  • 从未标记的网格点出发,将相邻(4邻/8邻)且高程差异在阈值内的点纳入当前簇
  • 簇生长完成后,仅保留点数≥4的有效簇,小簇可选择丢弃或合并到邻近相似簇

关键实现步骤

    1. 预处理:将二维高程数组转为可遍历网格,初始化访问标记矩阵
    1. 遍历未访问点,启动区域生长:
    • 初始化簇队列,加入当前点并标记为已访问
    • 循环取出队列内的点,检查所有邻域点:未访问且高程差≤阈值则加入队列
    • 生长结束后,若簇大小≥4则保存为有效簇
    1. 后处理:对小簇(<4)可合并到高程最接近的邻近有效簇,或直接丢弃

Python代码示例

基于numpy实现基础版区域生长,输入为二维高程数组elev_grid:

import numpy as np

def region_growing_clustering(elev_grid, elevation_threshold=0.5, min_cluster_size=4, connectivity=4):
    rows, cols = elev_grid.shape
    visited = np.zeros_like(elev_grid, dtype=bool)
    clusters = []
    cluster_id = 0

    # 定义邻域方向(4邻/8邻)
    directions = [(-1,0), (1,0), (0,-1), (0,1)] if connectivity ==4 else \
                 [(-1,-1), (-1,0), (-1,1), (0,-1), (0,1), (1,-1), (1,0), (1,1)]

    for i in range(rows):
        for j in range(cols):
            if not visited[i,j]:
                current_cluster = []
                queue = [(i,j)]
                visited[i,j] = True
                base_elev = elev_grid[i,j]

                while queue:
                    x, y = queue.pop(0)
                    current_cluster.append((x,y))

                    # 遍历邻域点
                    for dx, dy in directions:
                        nx, ny = x + dx, y + dy
                        if 0 <= nx < rows and 0 <= ny < cols:
                            if not visited[nx, ny] and abs(elev_grid[nx, ny] - base_elev) <= elevation_threshold:
                                visited[nx, ny] = True
                                queue.append((nx, ny))

                # 保留符合大小要求的簇
                if len(current_cluster) >= min_cluster_size:
                    clusters.append({
                        'id': cluster_id,
                        'points': current_cluster,
                        'mean_elev': np.mean([elev_grid[x,y] for x,y in current_cluster])
                    })
                    cluster_id += 1

    # 生成簇标记矩阵
    cluster_map = np.zeros_like(elev_grid, dtype=int)
    for cluster in clusters:
        for x,y in cluster['points']:
            cluster_map[x,y] = cluster['id'] + 1  # 0代表未聚类区域

    return cluster_map, clusters

参数说明

  • elevation_threshold:相邻点可纳入同一簇的最大高程差,需根据你的数据范围调整(如高程单位为米,可设1-5)
  • min_cluster_size:固定设为4,匹配你的需求
  • connectivity:选择4邻域(上下左右)或8邻域(含对角线),根据你对“连续”的定义调整

可视化实现

用matplotlib展示聚类结果:

import matplotlib.pyplot as plt

# 假设elev_grid是你的二维高程数据
cluster_map, clusters = region_growing_clustering(elev_grid, elevation_threshold=1.0)

plt.figure(figsize=(12,6))
plt.subplot(121)
plt.imshow(elev_grid, cmap='terrain')
plt.title('原始高程网格')
plt.colorbar(label='高程')

plt.subplot(122)
plt.imshow(cluster_map, cmap='tab20')
plt.title('聚类结果(连续簇,最小大小≥4)')
plt.colorbar(label='簇ID')
plt.show()

优化建议

  • 若高程数据噪声大,先做高斯平滑处理,减少无效小簇
  • 对小簇(<4),可遍历其邻近有效簇,合并到高程最接近的簇,提升结果完整性
  • 大数据量场景下,改用OpenCV的cv2.floodFill函数,效率更高

内容的提问来源于stack exchange,提问作者tincan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 22:59:59