基于TF-IDF的NLP评论聚类后可视化实现技术求助
嘿,你已经把聚类的核心工作搞定了,接下来的可视化其实不难,我给你整理了几个实用的图表实现,直接把这些代码片段加到你的现有代码里就能跑起来~
1. 先调整你的聚类函数(获取可视化所需变量)
原来的cluster_sentences函数只返回了聚类结果,咱们稍微改一下,把TF-IDF矩阵、KMeans模型和向量器也返回,方便后续可视化:
def cluster_sentences(sentences, nb_of_clusters=5): tfidf_vectorizer = TfidfVectorizer(tokenizer=word_tokenizer, stop_words=stopwords.words('english'), max_df=0.95,min_df=0.05, lowercase=True) tfidf_matrix = tfidf_vectorizer.fit_transform(sentences) kmeans = KMeans(n_clusters=nb_of_clusters) kmeans.fit(tfidf_matrix) clusters = collections.defaultdict(list) for i, label in enumerate(kmeans.labels_): clusters[label].append(i) # 返回额外变量供可视化使用 return dict(clusters), tfidf_matrix, kmeans, tfidf_vectorizer
2. 聚类数量分布直方图
这个图能帮你快速摸清各个聚类的规模,看看哪个聚类的评论最多、哪个最少:
if __name__ == "__main__": sentences = data.Comment nclusters= 20 # 调用修改后的函数,获取所有需要的变量 clusters, tfidf_matrix, kmeans, tfidf_vectorizer = cluster_sentences(sentences, nclusters) # 统计每个聚类的样本数量 cluster_sizes = [len(clusters[cluster]) for cluster in range(nclusters)] # 绘制直方图 plt.figure(figsize=(12, 6)) plt.bar(range(nclusters), cluster_sizes, color='#4287f5') plt.title('Cluster Size Distribution', fontsize=14) plt.xlabel('Cluster ID', fontsize=12) plt.ylabel('Number of Comments', fontsize=12) plt.xticks(range(nclusters)) # 显示所有聚类ID plt.grid(axis='y', linestyle='--', alpha=0.7) plt.show()
3. 聚类散点图(PCA降维)
TF-IDF是高维特征,没法直接画散点图,咱们用PCA把它压缩到2维,就能直观看到不同聚类的分布情况,判断聚类的区分度:
from sklearn.decomposition import PCA # 把高维TF-IDF矩阵降维到2维 pca = PCA(n_components=2) tfidf_2d = pca.fit_transform(tfidf_matrix.toarray()) # 绘制散点图,每个聚类用不同颜色 plt.figure(figsize=(12, 8)) for cluster_id in range(nclusters): # 获取当前聚类所有样本的2维坐标 cluster_points = tfidf_2d[[idx for idx in clusters[cluster_id]]] plt.scatter(cluster_points[:, 0], cluster_points[:, 1], label=f'Cluster {cluster_id}', alpha=0.6) plt.title('Clusters in PCA-reduced TF-IDF Space', fontsize=14) plt.xlabel('PCA Component 1', fontsize=12) plt.ylabel('PCA Component 2', fontsize=12) plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left') # 把图例放到图外,避免遮挡 plt.grid(linestyle='--', alpha=0.7) plt.show()
4. 可选:每个聚类的关键词展示(帮你理解聚类主题)
为了搞清楚每个聚类到底讲什么,你可以输出每个聚类的Top N关键词(基于TF-IDF权重):
def get_top_cluster_keywords(tfidf_vectorizer, kmeans_model, top_n=5): cluster_centers = kmeans_model.cluster_centers_ feature_words = tfidf_vectorizer.get_feature_names_out() cluster_keywords = {} for cluster_idx, center in enumerate(cluster_centers): # 按权重从高到低排序,取前top_n个词 top_word_indices = center.argsort()[-top_n:][::-1] cluster_keywords[cluster_idx] = [feature_words[idx] for idx in top_word_indices] return cluster_keywords # 获取每个聚类的Top5关键词 top_keywords = get_top_cluster_keywords(tfidf_vectorizer, kmeans, top_n=5) print("=== 每个聚类的Top关键词 ===") pprint(top_keywords)
比如你给出的Cluster 6,输出的关键词应该是['RIH', 'DP', 'std', 'liner', 'hole'],和示例里的评论内容完全对应,能快速帮你定位聚类主题。
内容的提问来源于stack exchange,提问作者Chym123
相关产品推荐
相关产品推荐

