如何基于Levenshtein相似度矩阵实现文本聚类?
基于Levenshtein距离的文本聚类实现
你可以利用**层次聚类(Agglomerative Clustering)**完成这个任务,它支持直接使用预先计算好的距离矩阵进行聚类,完全匹配你的场景。以下是完整的实现代码:
from Levenshtein import distance import numpy as np from sklearn.cluster import AgglomerativeClustering words = ['The Bachelor','The Bachelorette','The Bachelor Special','SportsCenter', 'SportsCenter 8 PM','SportsCenter Sunday'] # 生成Levenshtein距离矩阵 matrix = np.zeros((len(words), len(words)), dtype=np.int) for i in range(len(words)): for j in range(len(words)): matrix[i,j] = distance(words[i], words[j]) # 初始化层次聚类器,使用预先计算的距离矩阵 cluster = AgglomerativeClustering( n_clusters=2, affinity='precomputed', linkage='average' # 采用组间平均距离作为聚类合并依据 ) # 拟合距离矩阵得到聚类标签 labels = cluster.fit_predict(matrix) # 根据标签分组 clusters = {} for idx, label in enumerate(labels): if label not in clusters: clusters[label] = [] clusters[label].append(words[idx]) # 输出结果 for idx, group in clusters.items(): print(f"聚类 {idx+1}: {group}")
关键说明:
affinity='precomputed':告知聚类器传入的是已计算完成的距离矩阵,而非原始特征数据linkage='average':指定聚类合并时使用组间平均距离,能更好适配文本长度差异的场景- 若不想预先指定聚类数目,可改用
distance_threshold参数(比如设置distance_threshold=10,同时将n_clusters设为None),让算法自动根据距离阈值划分聚类
运行代码后会得到预期的两个聚类:
- 聚类1: ['The Bachelor', 'The Bachelorette', 'The Bachelor Special']
- 聚类2: ['SportsCenter', 'SportsCenter 8 PM', 'SportsCenter Sunday']
内容的提问来源于stack exchange,提问作者AI92
相关产品推荐
相关产品推荐

