You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言行业数据K-means聚类后添加标签方法求助

解决R语言聚类后添加观测标签的问题

Hey there! Let's walk through this step by step to get you the clustered dataset you want, similar to the US Crime example.

步骤1:整理数值型聚类数据集

既然K-means只支持数值变量,我们先确保用来聚类的是纯数值型数据(你已经移除了Industry因子,这一步没问题),同时标准化数据——这对K-means非常重要,因为它对变量尺度敏感:

library(tidyverse)

# 假设你的原始数据集名为df
# 筛选数值型变量并标准化
numeric_data <- df %>%
  select_if(is.numeric) %>%
  mutate_all(scale)

步骤2:确定最佳聚类数k

用肘部法则帮你选合适的聚类数量:

# 计算不同k值的组内平方和
wss_values <- map_dbl(1:10, function(k) {
  kmeans(numeric_data, centers = k, nstart = 20)$tot.withinss
})

# 绘制肘部图
tibble(clusters = 1:10, wss = wss_values) %>%
  ggplot(aes(x = clusters, y = wss)) +
  geom_line(linewidth = 1) +
  geom_point(size = 2) +
  labs(title = "Elbow Method for Optimal k", 
       x = "Number of Clusters", 
       y = "Within-Cluster Sum of Squares")

观察图中“肘部”位置(也就是组内平方和下降速度突然变慢的点),那就是最适合你的k值。

步骤3:运行K-means并添加聚类标签

选好k值后,执行聚类并把标签合并回原始数据集(包括你之前保留的所有变量,比如Industry):

# 替换成你从肘部图里选的k值,比如k=3
optimal_k <- 3
kmeans_result <- kmeans(numeric_data, centers = optimal_k, nstart = 20)

# 把聚类标签添加到原始数据集
final_clustered_df <- df %>%
  mutate(cluster = as.factor(kmeans_result$cluster))

现在final_clustered_df的结构就和US Crime数据集类似了:每个观测都有自己的聚类标签,同时保留了原始的Industry信息和其他变量。

额外:探索聚类与行业的关联

既然你原本关注行业维度,聚类后可以快速查看不同聚类里的行业分布:

final_clustered_df %>%
  count(cluster, Industry) %>%
  ggplot(aes(x = cluster, y = n, fill = Industry)) +
  geom_bar(stat = "identity") +
  labs(title = "Industry Distribution Across Clusters", 
       x = "Cluster", 
       y = "Number of Observations")

小提示

  • nstart = 20是让K-means多次初始化,避免陷入局部最优解,建议保留这个参数。
  • 如果你的数据有缺失值,记得先用drop_na()或者填充方法处理,否则K-means会报错。

内容的提问来源于stack exchange,提问作者13ISPaulGeorge

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:52:28