AgglomerativeClustering的distance_threshold参数科学选择方法
生产环境选择AgglomerativeClustering的distance_threshold的科学方法
1. 树状图肘部法则
先绘制聚类的树状图(dendrogram),观察聚类合并的距离变化拐点:
- 用scipy的
linkage和dendrogram函数绘制完整的聚类合并过程 - 找到合并距离斜率突变的「肘部」位置,该位置对应的距离就是最优的distance_threshold取值,你样例里的肘部位置刚好就在1.5附近,和手动调试的结果一致。
2. 聚类评估指标遍历法
预设一组合理的distance_threshold候选区间,对每个阈值跑出聚类结果后,用无监督聚类评估指标打分,选得分最高的阈值即可,常用的两个指标:
- 轮廓系数:取值范围[-1,1],越接近1说明簇内相似度越高、簇间差异越大,适合中小数据集
- Calinski-Harabasz指数:数值越高聚类效果越好,计算速度远快于轮廓系数,适合大规模数据集
参考实现代码:
from sklearn.metrics import silhouette_score, calinski_harabasz_score import numpy as np # 生成候选阈值区间,可根据样本距离范围调整步长和上下限 threshold_candidates = np.arange(0.5, 2.5, 0.1) best_score = -1 best_threshold = 1.5 for threshold in threshold_candidates: ag = AgglomerativeClustering(n_clusters=None, distance_threshold=threshold, linkage='ward') labels = ag.fit_predict(X_dense) # 过滤只生成1个簇、或每个样本单独成簇的无效结果 if len(set(labels)) == 1 or len(set(labels)) == len(X_dense): continue # 此处用轮廓系数打分,也可替换为calinski_harabasz_score score = silhouette_score(X_dense, labels) if score > best_score: best_score = score best_threshold = threshold print(f"最优阈值: {best_threshold}, 对应轮廓系数: {best_score}")
3. 业务先验约束法
如果业务场景对聚类的簇数有合理预期范围,可以先筛选出能让聚类簇数落在预期区间的threshold候选集合,再用上述评估指标从候选集合里选最优,能大幅降低调参成本,也能保证聚类结果符合业务逻辑。
注意事项
- distance_threshold的取值和
linkage参数强绑定,更换linkage规则(比如从ward换成average)后需要重新选择阈值,不能直接沿用之前的取值 - 高维特征(比如示例里的TF-IDF向量)使用ward linkage时默认用欧氏距离,高维下欧氏距离区分度会下降,可以先做PCA降维后再聚类,阈值选择的稳定性会更高
- 如果用余弦距离作为度量,只能选择average、complete、single linkage,不能使用ward linkage
内容的提问来源于stack exchange,提问作者A3006
相关产品推荐
相关产品推荐

