You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言K-means聚类图添加标签及解读聚类的技术问询

Hey there! Let's tackle your two K-means clustering questions one by one—super common pain points when working with text clustering, so I get where you're coming from.


1. Adding Document Labels to Your Clustering Plot

Your original text() call was pointing to the wrong data source (you used m_norm instead of the PCA projection coordinates). Here's a corrected, cleaner approach to plot your PCA results and add document labels clearly:

# First, compute PCA and store the result for easy reuse
pca_result <- prcomp(m_norm)

# Plot the PCA points with cluster colors (pch=16 makes points more visible)
plot(
  pca_result$x, 
  col = cl$cluster, 
  pch = 16, 
  main = "K-Means Clustering (5 Clusters) - PCA Projection",
  xlab = "Principal Component 1",
  ylab = "Principal Component 2"
)

# Add document labels (row names) to each data point
text(
  x = pca_result$x[,1], 
  y = pca_result$x[,2], 
  labels = rownames(m_norm), 
  cex = 0.7,  # Shrink label size to avoid overlap
  pos = 3     # Place labels above points
)

If your document names are too long, you can shorten them with abbreviate(rownames(m_norm), minlength=5) in the labels argument to keep the plot readable. For even better interactivity (hover to see labels), try the plotly package:

library(plotly)
plot_ly(
  x = pca_result$x[,1], 
  y = pca_result$x[,2], 
  color = factor(cl$cluster), 
  text = rownames(m_norm), 
  type = "scatter", 
  mode = "markers"
)

2. Interpreting table(cl$cluster) and Cluster Content

Let's break this down into two parts:

What does table(cl$cluster) mean?

The output numbers are simply the count of documents in each cluster. For example, if you see:

1   2   3   4   5 
 85  92  78  80  65 

That means Cluster 1 has 85 documents, Cluster 2 has 92, and so on. The total should add up to your full dataset size (400+ documents).

How to interpret the actual content of each cluster?

To understand what each cluster is about, you need to look at the most representative terms (words) for each group. Here are two reliable methods:

Method 1: Extract top-weighted terms per cluster

K-means gives you cluster centers (cl$centers) which represent the average TF-IDF weight of each term in the cluster. We can sort these weights to find the most important words:

# Get the list of terms from your document-term matrix
terms <- colnames(dtm_tfxidf)

# Extract top 10 terms for each cluster
top_terms <- lapply(1:5, function(cluster_num) {
  # Get center weights for the cluster
  cluster_weights <- cl$centers[cluster_num, ]
  # Sort terms by weight (highest first)
  sorted_terms <- sort(cluster_weights, decreasing = TRUE)
  # Grab top 10
  head(sorted_terms, 10)
})

# Name the list for clarity
names(top_terms) <- paste("Cluster", 1:5)

# Print the results
top_terms

If Cluster 1's top terms are "climate", "emission", "policy", you can infer that cluster focuses on climate change policy documents.

Method 2: Inspect actual documents in a cluster

To validate your term-based conclusions, you can directly view which documents belong to a cluster. Assuming your original documents are stored in a data frame/vector called docs_data:

# Add cluster labels to your document dataset
docs_data$cluster <- cl$cluster

# View all documents in Cluster 1, for example
subset(docs_data, cluster == 1)

This lets you read sample documents and confirm the cluster's theme matches the top terms.


内容的提问来源于stack exchange,提问作者daveyjones

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:16:27