在R语言K-means聚类图添加标签及解读聚类的技术问询
Hey there! Let's tackle your two K-means clustering questions one by one—super common pain points when working with text clustering, so I get where you're coming from.
Your original text() call was pointing to the wrong data source (you used m_norm instead of the PCA projection coordinates). Here's a corrected, cleaner approach to plot your PCA results and add document labels clearly:
# First, compute PCA and store the result for easy reuse pca_result <- prcomp(m_norm) # Plot the PCA points with cluster colors (pch=16 makes points more visible) plot( pca_result$x, col = cl$cluster, pch = 16, main = "K-Means Clustering (5 Clusters) - PCA Projection", xlab = "Principal Component 1", ylab = "Principal Component 2" ) # Add document labels (row names) to each data point text( x = pca_result$x[,1], y = pca_result$x[,2], labels = rownames(m_norm), cex = 0.7, # Shrink label size to avoid overlap pos = 3 # Place labels above points )
If your document names are too long, you can shorten them with abbreviate(rownames(m_norm), minlength=5) in the labels argument to keep the plot readable. For even better interactivity (hover to see labels), try the plotly package:
library(plotly) plot_ly( x = pca_result$x[,1], y = pca_result$x[,2], color = factor(cl$cluster), text = rownames(m_norm), type = "scatter", mode = "markers" )
table(cl$cluster) and Cluster Content Let's break this down into two parts:
What does table(cl$cluster) mean?
The output numbers are simply the count of documents in each cluster. For example, if you see:
1 2 3 4 5 85 92 78 80 65
That means Cluster 1 has 85 documents, Cluster 2 has 92, and so on. The total should add up to your full dataset size (400+ documents).
How to interpret the actual content of each cluster?
To understand what each cluster is about, you need to look at the most representative terms (words) for each group. Here are two reliable methods:
Method 1: Extract top-weighted terms per cluster
K-means gives you cluster centers (cl$centers) which represent the average TF-IDF weight of each term in the cluster. We can sort these weights to find the most important words:
# Get the list of terms from your document-term matrix terms <- colnames(dtm_tfxidf) # Extract top 10 terms for each cluster top_terms <- lapply(1:5, function(cluster_num) { # Get center weights for the cluster cluster_weights <- cl$centers[cluster_num, ] # Sort terms by weight (highest first) sorted_terms <- sort(cluster_weights, decreasing = TRUE) # Grab top 10 head(sorted_terms, 10) }) # Name the list for clarity names(top_terms) <- paste("Cluster", 1:5) # Print the results top_terms
If Cluster 1's top terms are "climate", "emission", "policy", you can infer that cluster focuses on climate change policy documents.
Method 2: Inspect actual documents in a cluster
To validate your term-based conclusions, you can directly view which documents belong to a cluster. Assuming your original documents are stored in a data frame/vector called docs_data:
# Add cluster labels to your document dataset docs_data$cluster <- cl$cluster # View all documents in Cluster 1, for example subset(docs_data, cluster == 1)
This lets you read sample documents and confirm the cluster's theme matches the top terms.
内容的提问来源于stack exchange,提问作者daveyjones

