R语言层次聚类:如何提取簇对应行生成独立数据框?
Hey there! Let's get those cluster-specific data frames sorted out for you—you're already halfway there with your cluster assignments from cutree(). Here are two reliable, straightforward ways to split your original Complete_df into separate data frames for each cluster:
Method 1: Use split() (Recommended for Organization)
First, let's attach your cluster labels to the original data frame to make our operations explicit:
# Add cluster assignment as a new column in your data frame Complete_df$cluster <- cutree_hclust
Then use the split() function to create a list where each element is a data frame corresponding to one cluster:
# Split the data frame by cluster cluster_list <- split(Complete_df, f = Complete_df$cluster)
You can access individual cluster data frames using list indexing:
cluster_list[[1]]will give you all rows from cluster 1 (867 observations)cluster_list[["5"]]will pull up cluster 5's data (1135 observations)
This method keeps all your cluster data organized in one place, which is perfect for iterative analysis or applying functions across all clusters.
Method 2: Export to Separate Global Environment Objects
If you want each cluster as a standalone data frame in your global environment (e.g., cluster_1, cluster_2, etc.), you can use list2env() after renaming the list elements for clarity:
# Rename list elements to have consistent, readable names names(cluster_list) <- paste0("cluster_", names(cluster_list)) # Export each cluster data frame to the global environment list2env(cluster_list, envir = .GlobalEnv)
Now you can directly call cluster_1, cluster_8, etc., to work with each group's data without digging into a list.
Quick Validation Check
To confirm everything worked correctly, cross-reference the row count of any cluster data frame with your table(cutree_hclust) output:
# Verify cluster 1 has 867 rows nrow(cluster_list[[1]]) # Or nrow(cluster_1) if you used Method 2
内容的提问来源于stack exchange,提问作者BloopFloopy

