You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Levenshtein相似度矩阵实现文本聚类?

基于Levenshtein距离的文本聚类实现

你可以利用**层次聚类(Agglomerative Clustering)**完成这个任务,它支持直接使用预先计算好的距离矩阵进行聚类,完全匹配你的场景。以下是完整的实现代码:

from Levenshtein import distance
import numpy as np
from sklearn.cluster import AgglomerativeClustering

words = ['The Bachelor','The Bachelorette','The Bachelor Special','SportsCenter',
         'SportsCenter 8 PM','SportsCenter Sunday']

# 生成Levenshtein距离矩阵
matrix = np.zeros((len(words), len(words)), dtype=np.int)
for i in range(len(words)):
    for j in range(len(words)):
        matrix[i,j] = distance(words[i], words[j])

# 初始化层次聚类器,使用预先计算的距离矩阵
cluster = AgglomerativeClustering(
    n_clusters=2,
    affinity='precomputed',
    linkage='average'  # 采用组间平均距离作为聚类合并依据
)
# 拟合距离矩阵得到聚类标签
labels = cluster.fit_predict(matrix)

# 根据标签分组
clusters = {}
for idx, label in enumerate(labels):
    if label not in clusters:
        clusters[label] = []
    clusters[label].append(words[idx])

# 输出结果
for idx, group in clusters.items():
    print(f"聚类 {idx+1}: {group}")

关键说明:

  • affinity='precomputed':告知聚类器传入的是已计算完成的距离矩阵,而非原始特征数据
  • linkage='average':指定聚类合并时使用组间平均距离,能更好适配文本长度差异的场景
  • 若不想预先指定聚类数目,可改用distance_threshold参数(比如设置distance_threshold=10,同时将n_clusters设为None),让算法自动根据距离阈值划分聚类

运行代码后会得到预期的两个聚类:

  • 聚类1: ['The Bachelor', 'The Bachelorette', 'The Bachelor Special']
  • 聚类2: ['SportsCenter', 'SportsCenter 8 PM', 'SportsCenter Sunday']

内容的提问来源于stack exchange,提问作者AI92

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 04:01:14