在R中基于皮尔逊相关系数阈值手动实现样本聚类
R实现自定义阈值手动聚类方案
核心逻辑说明
你需要的聚类规则本质是无向图的连通分量查找:两个样本皮尔逊相关系数大于阈值则存在连接,存在连接路径的所有样本归为同一簇,不需要调用kmeans、层次聚类等现成聚类算法。
注意:你之前的代码用
c作为变量名容易和R内置的c()函数冲突,建议修改为cor_mat避免运行报错。
完整实现代码
步骤1:复现你已有的距离计算逻辑
# 计算样本间皮尔逊相关系数 cor_mat <- round(cor(t(df)), digits = 2) # 转换为距离矩阵(距离=1-相关系数) d <- as.dist(1 - cor_mat)
步骤2:设定阈值并构造邻接矩阵
# 自定义相关系数阈值,可按需修改 cor_threshold <- 0.7 # 构造邻接矩阵:值为TRUE代表两个样本满足同簇条件 adj_matrix <- cor_mat >= cor_threshold # 强制对角线为TRUE,样本自身默认属于同一簇 diag(adj_matrix) <- TRUE
步骤3:查找连通分量得到聚类编号
提供两种实现方式,按需选择:
方式1:无额外依赖,手动深度优先搜索实现
n_sample <- nrow(df) visited <- rep(FALSE, n_sample) cluster_id <- rep(NA, n_sample) current_cluster_num <- 1 for (i in 1:n_sample) { if (!visited[i]) { search_queue <- i visited[i] <- TRUE cluster_id[i] <- current_cluster_num # 深度优先搜索所有连通节点 while (length(search_queue) > 0) { current_node <- search_queue[1] search_queue <- search_queue[-1] # 找到所有和当前节点连通的未访问样本 neighbors <- which(adj_matrix[current_node, ] & !visited) visited[neighbors] <- TRUE cluster_id[neighbors] <- current_cluster_num search_queue <- c(search_queue, neighbors) } current_cluster_num <- current_cluster_num + 1 } }
方式2:使用igraph包简化计算(仅做图连通性计算,未调用聚类算法)
# 未装包先执行 install.packages("igraph") library(igraph) # 从邻接矩阵构造无向图 g <- graph_from_adjacency_matrix(adj_matrix, mode = "undirected", diag = FALSE) # 提取连通分量编号 cluster_id <- components(g)$membership
步骤4:将聚类编号写入原数据框
df$cluster_id <- cluster_id
示例结果匹配说明
对应你提供的距离矩阵,阈值设为0.7时聚类结果为:
- 簇1:U00、U01(相关系数0.95>0.7)
- 簇2:U02(和其他所有样本相关系数最高仅0.08<0.7)
- 簇3:U03、U04、U05、U06(两两相关系数均为1>0.7)
内容的提问来源于stack exchange,提问作者ncnc_2020
相关产品推荐
相关产品推荐

