R语言行业数据K-means聚类后添加标签方法求助
解决R语言聚类后添加观测标签的问题
Hey there! Let's walk through this step by step to get you the clustered dataset you want, similar to the US Crime example.
步骤1:整理数值型聚类数据集
既然K-means只支持数值变量,我们先确保用来聚类的是纯数值型数据(你已经移除了Industry因子,这一步没问题),同时标准化数据——这对K-means非常重要,因为它对变量尺度敏感:
library(tidyverse) # 假设你的原始数据集名为df # 筛选数值型变量并标准化 numeric_data <- df %>% select_if(is.numeric) %>% mutate_all(scale)
步骤2:确定最佳聚类数k
用肘部法则帮你选合适的聚类数量:
# 计算不同k值的组内平方和 wss_values <- map_dbl(1:10, function(k) { kmeans(numeric_data, centers = k, nstart = 20)$tot.withinss }) # 绘制肘部图 tibble(clusters = 1:10, wss = wss_values) %>% ggplot(aes(x = clusters, y = wss)) + geom_line(linewidth = 1) + geom_point(size = 2) + labs(title = "Elbow Method for Optimal k", x = "Number of Clusters", y = "Within-Cluster Sum of Squares")
观察图中“肘部”位置(也就是组内平方和下降速度突然变慢的点),那就是最适合你的k值。
步骤3:运行K-means并添加聚类标签
选好k值后,执行聚类并把标签合并回原始数据集(包括你之前保留的所有变量,比如Industry):
# 替换成你从肘部图里选的k值,比如k=3 optimal_k <- 3 kmeans_result <- kmeans(numeric_data, centers = optimal_k, nstart = 20) # 把聚类标签添加到原始数据集 final_clustered_df <- df %>% mutate(cluster = as.factor(kmeans_result$cluster))
现在final_clustered_df的结构就和US Crime数据集类似了:每个观测都有自己的聚类标签,同时保留了原始的Industry信息和其他变量。
额外:探索聚类与行业的关联
既然你原本关注行业维度,聚类后可以快速查看不同聚类里的行业分布:
final_clustered_df %>% count(cluster, Industry) %>% ggplot(aes(x = cluster, y = n, fill = Industry)) + geom_bar(stat = "identity") + labs(title = "Industry Distribution Across Clusters", x = "Cluster", y = "Number of Observations")
小提示
nstart = 20是让K-means多次初始化,避免陷入局部最优解,建议保留这个参数。- 如果你的数据有缺失值,记得先用
drop_na()或者填充方法处理,否则K-means会报错。
内容的提问来源于stack exchange,提问作者13ISPaulGeorge
相关产品推荐
相关产品推荐

