如何基于聚类点数量快速过滤LiDAR点数据?
问题描述
使用DBSCAN对大量LiDAR点进行聚类以优化植被点分类,聚类后生成cluster_ID列标识每个点所属聚类组。需要移除点数低于指定阈值的聚类,当前用for循环+if语句处理7万余点、约600类的数据耗时约2分钟(仅为实际数据子集),希望优化提速。
现有可行但效率低下的代码示例:
rm(list=ls()) # clear env # make empty list test_data <- c() # add data to cluster by test_data$cluster_ID <- as.integer(runif(10,min = 1, max = 5)) # convert to df test_data <- as.data.frame(test_data) # summarize by cluster sum_test <- test_data %>% group_by(cluster_ID) %>% summarise(pt_count = sum(cluster_ID)) # convert from sum to num points sum_test$pt_count <- sum_test$pt_count/sum_test$cluster_ID # remove all rows < 3 points in "cluster_ID" for (c_id in unique(test_data$cluster_ID)) { # get index of unique c_id idx <- which(test_data$cluster_ID == c_id) # filter out length < 3 and remove from test_data if (length(idx) < 3) { test_data <- test_data[-c(idx)] } }
核心需求:基于summarise的输出,不使用循环筛选出test_data中点数低于阈值的行(因test_data与sum_test长度不等,无法直接匹配)。
优化方案
1. 修正聚类点数统计逻辑
原代码中用sum(cluster_ID)再除以cluster_ID统计点数的方式错误且冗余,直接用n()即可获取每个聚类的点数:
sum_test <- test_data %>% group_by(cluster_ID) %>% summarise(pt_count = n()) # 直接统计每个聚类的点数
2. 高效筛选符合阈值的聚类并过滤原数据
无需循环,通过向量匹配或半连接即可快速完成过滤,两种方案任选:
方案一:使用%in%向量匹配
threshold <- 3 # 筛选出点数≥阈值的cluster_ID valid_clusters <- sum_test %>% filter(pt_count >= threshold) %>% pull(cluster_ID) # 过滤原数据,仅保留有效聚类的点 test_data_filtered <- test_data %>% filter(cluster_ID %in% valid_clusters)
方案二:使用semi_join半连接
半连接会直接保留test_data中与sum_test筛选结果匹配的行,无需额外提取ID列表:
threshold <- 3 test_data_filtered <- test_data %>% semi_join( sum_test %>% filter(pt_count >= threshold), by = "cluster_ID" )
效率说明
- 上述方案均基于dplyr的向量化操作,避免了循环中反复修改数据框的开销,处理大数据量时速度会提升数十倍甚至上百倍。
- 对于7万级别的数据,处理时间可压缩至秒级。
内容的提问来源于stack exchange,提问作者nola
相关产品推荐
相关产品推荐

