如何在R中将无监督层次聚类结果合并至原始数据
Got it, let's walk through how to seamlessly add your cluster groups to your original dataset so you can move forward with spatial mapping. Here's a clear breakdown based on your existing code:
Recap of Your Clustering Workflow
First, let's confirm your existing code is properly formatted (you already have the core clustering steps down!):
# Perform hierarchical clustering using Ward's minimum variance method hc5 <- hclust(d, method = "ward.D2") # Cut the dendrogram into 6 distinct clusters sub_grp <- cutree(hc5, k = 6) # Check the number of observations in each cluster table(sub_grp)
Merge Cluster Labels with Your Original Dataset
The sub_grp vector you created is ordered to match the rows in your distance matrix d. As long as the row order in your original data frame matches the order used to compute d, adding the cluster column is straightforward.
Base R Approach
Replace original_data with the actual name of your 120k-row data frame:
# Add cluster group as a new column original_data$cluster <- sub_grp
Tidyverse Approach (if you use dplyr)
If you prefer the tidy syntax, this works too:
library(dplyr) original_data <- original_data %>% mutate(cluster = sub_grp)
Verify the Merge Worked
Quick checks to ensure everything aligns correctly:
# View the first few rows with the new cluster column head(original_data) # Confirm cluster counts match your earlier table output table(original_data$cluster)
Quick Tip for Spatial Mapping
Now that your data has cluster labels, make sure you have your spatial coordinate columns (like lat/long) intact. For plotting, you can use packages like ggplot2 (with geom_point(aes(color = factor(cluster)))) or sf for more advanced spatial layers.
内容的提问来源于stack exchange,提问作者Z3804

