You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scikit从K-Means聚类中提取高频关键词?

提取并可视化聚类高频词的方法

一、提取每个聚类的高频词

方法1:基于KMeans聚类中心的权重排序

KMeans的cluster_centers_属性存储了每个聚类在TF-IDF特征空间的中心向量,每个元素对应词汇的TF-IDF权重,通过排序即可提取权重最高的N个词:

import numpy as np

# 获取词汇表,建立词与特征索引的映射
terms = vectorizer.get_feature_names_out()
true_k = 9
top_n = 10  # 每个聚类提取前10个高频词

# 遍历所有聚类,提取并输出高频词
for i in range(true_k):
    # 对当前聚类的中心向量排序,取权重最高的top_n个词的索引
    center_indices = model.cluster_centers_[i].argsort()[-top_n:][::-1]
    # 转换为对应的词汇
    top_terms = [terms[idx] for idx in center_indices]
    print(f"聚类 {i} 的Top {top_n} 高频词:")
    print(", ".join(top_terms))
    print("-"*50)

方法2:按聚类分组计算平均TF-IDF权重

对每个聚类内的所有文档TF-IDF向量取平均值,再排序提取高频词,更贴合组内实际的词频特征:

# 将TF-IDF稀疏矩阵转换为DataFrame,行对应文档,列对应词汇
tfidf_df = pd.DataFrame(X.toarray(), columns=terms)
# 加入聚类标签列
tfidf_df['Cluster'] = data_cl['Cluster']

# 按聚类分组,计算每个词汇的平均TF-IDF值
cluster_term_means = tfidf_df.groupby('Cluster').mean()

top_n = 10
# 遍历所有聚类,提取并输出高频词
for cluster in range(true_k):
    # 对当前聚类的平均TF-IDF值降序排序,取前top_n个
    sorted_terms = cluster_term_means.loc[cluster].sort_values(ascending=False).head(top_n)
    print(f"聚类 {cluster} 的Top {top_n} 高频词:")
    print(", ".join(sorted_terms.index))
    print("-"*50)

二、可视化高频词

1. 条形图可视化

用matplotlib绘制条形图,直观展示每个聚类高频词的权重差异:

import matplotlib.pyplot as plt

top_n = 10
# 遍历所有聚类
for cluster in range(true_k):
    sorted_terms = cluster_term_means.loc[cluster].sort_values(ascending=False).head(top_n)
    # 绘制横向条形图
    plt.figure(figsize=(10, 6))
    plt.barh(sorted_terms.index, sorted_terms.values, color='skyblue')
    plt.gca().invert_yaxis()  # 让权重高的词显示在上方
    plt.title(f"聚类 {cluster} 的Top {top_n} 高频词")
    plt.xlabel("平均TF-IDF权重")
    plt.ylabel("词汇")
    plt.tight_layout()
    plt.show()

2. 词云可视化

使用wordcloud库生成词云,更直观呈现高频词的重要性:
首先安装依赖库:

pip install wordcloud

然后运行生成词云的代码:

from wordcloud import WordCloud

# 遍历所有聚类
for cluster in range(true_k):
    # 拼接当前聚类的所有文本内容
    cluster_texts = data_cl[data_cl['Cluster'] == cluster]['TEXT'].str.cat(sep=' ')
    # 生成词云(自动过滤英文停用词)
    wordcloud = WordCloud(stop_words='english', background_color='white', width=800, height=600).generate(cluster_texts)
    # 展示词云
    plt.figure(figsize=(10, 8))
    plt.imshow(wordcloud, interpolation='bilinear')
    plt.axis('off')
    plt.title(f"聚类 {cluster} 词云")
    plt.show()

内容的提问来源于stack exchange,提问作者Mbando

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 17:27:30