You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于TF-IDF的NLP评论聚类后可视化实现技术求助

嘿,你已经把聚类的核心工作搞定了,接下来的可视化其实不难,我给你整理了几个实用的图表实现,直接把这些代码片段加到你的现有代码里就能跑起来~


1. 先调整你的聚类函数(获取可视化所需变量)

原来的cluster_sentences函数只返回了聚类结果,咱们稍微改一下,把TF-IDF矩阵、KMeans模型和向量器也返回,方便后续可视化:

def cluster_sentences(sentences, nb_of_clusters=5):
    tfidf_vectorizer = TfidfVectorizer(tokenizer=word_tokenizer, stop_words=stopwords.words('english'),
                                       max_df=0.95,min_df=0.05, lowercase=True)
    tfidf_matrix = tfidf_vectorizer.fit_transform(sentences)
    kmeans = KMeans(n_clusters=nb_of_clusters)
    kmeans.fit(tfidf_matrix)
    clusters = collections.defaultdict(list)
    for i, label in enumerate(kmeans.labels_):
        clusters[label].append(i)
    # 返回额外变量供可视化使用
    return dict(clusters), tfidf_matrix, kmeans, tfidf_vectorizer

2. 聚类数量分布直方图

这个图能帮你快速摸清各个聚类的规模,看看哪个聚类的评论最多、哪个最少:

if __name__ == "__main__":
    sentences = data.Comment
    nclusters= 20
    # 调用修改后的函数,获取所有需要的变量
    clusters, tfidf_matrix, kmeans, tfidf_vectorizer = cluster_sentences(sentences, nclusters)
    
    # 统计每个聚类的样本数量
    cluster_sizes = [len(clusters[cluster]) for cluster in range(nclusters)]
    
    # 绘制直方图
    plt.figure(figsize=(12, 6))
    plt.bar(range(nclusters), cluster_sizes, color='#4287f5')
    plt.title('Cluster Size Distribution', fontsize=14)
    plt.xlabel('Cluster ID', fontsize=12)
    plt.ylabel('Number of Comments', fontsize=12)
    plt.xticks(range(nclusters))  # 显示所有聚类ID
    plt.grid(axis='y', linestyle='--', alpha=0.7)
    plt.show()

3. 聚类散点图(PCA降维)

TF-IDF是高维特征,没法直接画散点图,咱们用PCA把它压缩到2维,就能直观看到不同聚类的分布情况,判断聚类的区分度:

from sklearn.decomposition import PCA

# 把高维TF-IDF矩阵降维到2维
pca = PCA(n_components=2)
tfidf_2d = pca.fit_transform(tfidf_matrix.toarray())

# 绘制散点图,每个聚类用不同颜色
plt.figure(figsize=(12, 8))
for cluster_id in range(nclusters):
    # 获取当前聚类所有样本的2维坐标
    cluster_points = tfidf_2d[[idx for idx in clusters[cluster_id]]]
    plt.scatter(cluster_points[:, 0], cluster_points[:, 1], label=f'Cluster {cluster_id}', alpha=0.6)

plt.title('Clusters in PCA-reduced TF-IDF Space', fontsize=14)
plt.xlabel('PCA Component 1', fontsize=12)
plt.ylabel('PCA Component 2', fontsize=12)
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')  # 把图例放到图外,避免遮挡
plt.grid(linestyle='--', alpha=0.7)
plt.show()

4. 可选:每个聚类的关键词展示(帮你理解聚类主题)

为了搞清楚每个聚类到底讲什么,你可以输出每个聚类的Top N关键词(基于TF-IDF权重):

def get_top_cluster_keywords(tfidf_vectorizer, kmeans_model, top_n=5):
    cluster_centers = kmeans_model.cluster_centers_
    feature_words = tfidf_vectorizer.get_feature_names_out()
    
    cluster_keywords = {}
    for cluster_idx, center in enumerate(cluster_centers):
        # 按权重从高到低排序,取前top_n个词
        top_word_indices = center.argsort()[-top_n:][::-1]
        cluster_keywords[cluster_idx] = [feature_words[idx] for idx in top_word_indices]
    return cluster_keywords

# 获取每个聚类的Top5关键词
top_keywords = get_top_cluster_keywords(tfidf_vectorizer, kmeans, top_n=5)
print("=== 每个聚类的Top关键词 ===")
pprint(top_keywords)

比如你给出的Cluster 6,输出的关键词应该是['RIH', 'DP', 'std', 'liner', 'hole'],和示例里的评论内容完全对应,能快速帮你定位聚类主题。


内容的提问来源于stack exchange,提问作者Chym123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:11:03