使用pheatmap绘制热图时聚类结果不符合预期,请求协助排查
解决pheatmap聚类结果不符合预期的问题
使用pheatmap绘制数据热图时,聚类结果不符合预期:cluster1与cluster2的聚类距离比cluster3更近,但实际cluster2应与前两者差异更大、距离更远。注意:不希望对数据进行缩放。
当前热图结果:
当前使用的脚本
pheatmap(data, show_colnames = T, show_rownames = T,fontsize=12,angle_col = 45, color=colorRampPalette(c('navy', "white", "#EA0764"))(100))
数据
> dput(data) structure(list(BRCA = c(0.230749533, 0.070357349, -0.239713974, 0.703365372, 0.327406476, 0.788362186, -0.145366946, 0.193414478, -0.427479686, -0.548639578, -0.227693913, 0.007319619, -0.282273491, -0.127434584, -0.465189842, 0.075561033), COAD = c(0.332890058, 0.09527495, -0.178301154, 0.924092152, 0.612677767, 1.221923426, -0.276093973, 0.958147584, 0.379735117, 0.029502217, -0.054424396, 0.300349746, 0.650179747, 0.155807135, 0.028132132, 0.378762531 ), LUAD = c(0.829893225, 0.2606768, -0.1821214, 0.692721414, 0.688720644, 1.017873137, 0.056591306, 0.289392746, 0.005723013, -0.388952621, 0.256912977, 0.186641748, 0.017661571, -0.270885042, -0.16046479, -0.097848349), LUSC = c(1.159170792, 0.533298568, 0.01576386, 0.901518727, 1.14831359, 1.40438067, 0.150671917, 0.627636011, -0.221484602, -0.621483798, 0.104588174, 0.377637951, 0.427781729, -0.444719169, -0.26419915, 0.042390119), STAD = c(0.494758767, 0.125857219, -0.016071377, 0.203326756, 0.784281834, 1.178387864, 0.159516768, 0.711241935, -0.03263392, -0.30666277, -0.056459557, -0.202547246, 0.182941332, -0.139104068, -0.10655886, 0.262286026 ), THCA = c(-0.179233856, -0.035498075, -0.634293378, -0.018114387, -0.029922798, 0.21657491, 0.040044405, -0.085685179, -0.33303136, 0.206747289, -0.112778263, 0.115131275, -0.217347418, -0.178894761, -0.36714963, -0.186071557), UCEC = c(0.228783922, 0.087425611, -0.126598752, 0.985393743, 0.203059279, 0.962485138, -0.050525831, 0.097276188, -0.386458244, -0.47970207, -0.542838744, 0.090891178, 0.146372654, -0.148096616, -0.404956997, -0.099997733), ESCA = c(0.515480135, -0.128144806, -0.001233162, 0.375547256, 0.868288817, 0.846731963, -0.09909923, 0.215571998, -0.327098337, -0.686121329, -0.332182132, 0.081704439, 0.022641456, -0.021981116, -0.51026229, 0.33227835 ), KIRC = c(-0.059448938, 0.061253477, -0.605107439, 0.135114673, 0.286596096, -0.085087897, 0.125670677, 0.394701809, 0.199182749, 0.400641612, 0.410932558, -0.159331938, 0.072832951, 0.281035001, 0.391288988, 0.066404822), PRAD = c(0.128646849, 0.009521855, 0.126533666, 0.542727651, 0.000694083, 0.32625358, -0.131356051, 0.087581413, -0.264907468, -0.201647563, -0.088937601, -0.110720233, 0.343355006, 0.205884841, 0.235804542, -0.049704744), LIHC = c(0.581419719, 0.121962843, -0.661213563, 0.219828076, 0.430757101, 0.935526611, 0.090115912, 1.00110355, -0.368156896, -0.264853361, -0.249296333, -0.216131716, -0.161492687, -0.096794496, 0.091178754, 0.353416867 ), GBM = c(-0.327762529, -0.100512851, -1.170250005, 0.973784849, 0.22308926, 0.605125243, -0.170939552, -0.213415522, -0.548838235, -0.236635581, -0.302241958, 0.09930227, 0.628040053, 0.11204205, -0.275029974, 0.540026962), BLCA = c(0.447543962, 0.296830523, -0.242805848, 1.077955557, 0.433840673, 0.78600795, -0.150633173, -0.11054011, -0.552050335, -0.27309979, 0.05763108, -0.30845115, 0.073324125, -0.396794707, -0.380437894, 0.130825992)), class = "data.frame", row.names = c("Gene1", "Gene2", "Gene3", "Gene4", "Gene5", "Gene6", "Gene7", "Gene8", "Gene9", "Gene10", "Gene11", "Gene12", "Gene13", "Gene14", "Gene15", "Gene16"))
调整方案
聚类结果不符合预期的核心原因是默认的距离度量和聚类算法不一定适配你的数据分布,可通过以下参数调整:
修改距离计算方法
默认使用欧氏距离,可尝试曼哈顿距离、皮尔逊相关系数等,针对列聚类调整用distance_col,行聚类用distance_row:# 示例:列聚类使用曼哈顿距离 pheatmap(data, show_colnames = T, show_rownames = T,fontsize=12,angle_col = 45, color=colorRampPalette(c('navy', "white", "#EA0764"))(100), distance_col = "manhattan")更换聚类算法
默认使用ward.D2,可尝试complete、average等方法,调整参数clustering_method:# 示例:使用complete聚类方法 pheatmap(data, show_colnames = T, show_rownames = T,fontsize=12,angle_col = 45, color=colorRampPalette(c('navy', "white", "#EA0764"))(100), clustering_method = "complete")组合调整
同时修改距离和聚类方法,找到最贴合预期的组合,比如:pheatmap(data, show_colnames = T, show_rownames = T,fontsize=12,angle_col = 45, color=colorRampPalette(c('navy', "white", "#EA0764"))(100), distance_col = "correlation", # 用相关系数衡量列间相似性 clustering_method = "average")验证聚类逻辑
可以先单独计算样本间的距离矩阵,确认距离关系是否符合预期:# 计算列间距离矩阵(以曼哈顿距离为例) dist_matrix <- dist(t(data), method = "manhattan") print(as.matrix(dist_matrix))通过距离矩阵可以直观看到样本间的实际距离,再选择合适的聚类方法。
内容的提问来源于stack exchange,提问作者zahra abdi
相关产品推荐
相关产品推荐

