如何用Scikit从K-Means聚类中提取高频关键词?
提取并可视化聚类高频词的方法
一、提取每个聚类的高频词
方法1:基于KMeans聚类中心的权重排序
KMeans的cluster_centers_属性存储了每个聚类在TF-IDF特征空间的中心向量,每个元素对应词汇的TF-IDF权重,通过排序即可提取权重最高的N个词:
import numpy as np # 获取词汇表,建立词与特征索引的映射 terms = vectorizer.get_feature_names_out() true_k = 9 top_n = 10 # 每个聚类提取前10个高频词 # 遍历所有聚类,提取并输出高频词 for i in range(true_k): # 对当前聚类的中心向量排序,取权重最高的top_n个词的索引 center_indices = model.cluster_centers_[i].argsort()[-top_n:][::-1] # 转换为对应的词汇 top_terms = [terms[idx] for idx in center_indices] print(f"聚类 {i} 的Top {top_n} 高频词:") print(", ".join(top_terms)) print("-"*50)
方法2:按聚类分组计算平均TF-IDF权重
对每个聚类内的所有文档TF-IDF向量取平均值,再排序提取高频词,更贴合组内实际的词频特征:
# 将TF-IDF稀疏矩阵转换为DataFrame,行对应文档,列对应词汇 tfidf_df = pd.DataFrame(X.toarray(), columns=terms) # 加入聚类标签列 tfidf_df['Cluster'] = data_cl['Cluster'] # 按聚类分组,计算每个词汇的平均TF-IDF值 cluster_term_means = tfidf_df.groupby('Cluster').mean() top_n = 10 # 遍历所有聚类,提取并输出高频词 for cluster in range(true_k): # 对当前聚类的平均TF-IDF值降序排序,取前top_n个 sorted_terms = cluster_term_means.loc[cluster].sort_values(ascending=False).head(top_n) print(f"聚类 {cluster} 的Top {top_n} 高频词:") print(", ".join(sorted_terms.index)) print("-"*50)
二、可视化高频词
1. 条形图可视化
用matplotlib绘制条形图,直观展示每个聚类高频词的权重差异:
import matplotlib.pyplot as plt top_n = 10 # 遍历所有聚类 for cluster in range(true_k): sorted_terms = cluster_term_means.loc[cluster].sort_values(ascending=False).head(top_n) # 绘制横向条形图 plt.figure(figsize=(10, 6)) plt.barh(sorted_terms.index, sorted_terms.values, color='skyblue') plt.gca().invert_yaxis() # 让权重高的词显示在上方 plt.title(f"聚类 {cluster} 的Top {top_n} 高频词") plt.xlabel("平均TF-IDF权重") plt.ylabel("词汇") plt.tight_layout() plt.show()
2. 词云可视化
使用wordcloud库生成词云,更直观呈现高频词的重要性:
首先安装依赖库:
pip install wordcloud
然后运行生成词云的代码:
from wordcloud import WordCloud # 遍历所有聚类 for cluster in range(true_k): # 拼接当前聚类的所有文本内容 cluster_texts = data_cl[data_cl['Cluster'] == cluster]['TEXT'].str.cat(sep=' ') # 生成词云(自动过滤英文停用词) wordcloud = WordCloud(stop_words='english', background_color='white', width=800, height=600).generate(cluster_texts) # 展示词云 plt.figure(figsize=(10, 8)) plt.imshow(wordcloud, interpolation='bilinear') plt.axis('off') plt.title(f"聚类 {cluster} 词云") plt.show()
内容的提问来源于stack exchange,提问作者Mbando
相关产品推荐
相关产品推荐

