You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中实现Kmeans聚类散点图的点名悬停显示

解决Kmeans聚类散点图悬停显示数据点名称的问题

核心问题是fviz_cluster在指定geom="point"时,未将数据点名称纳入ggplot的映射属性,导致ggplotly无法读取该信息。以下是两种可行的解决方法:

方法一:手动基于PCA结果构建绘图数据

这种方法更灵活,完全自定义绘图元素:

# 加载包
library(factoextra)
library(plotly)
library(stats)
library(ggplot2)

# 读取数据(假设CSV第一列为数据点名称,设为行名)
df <- read.csv("你的测试数据.csv", row.names = 1)

# 执行Kmeans聚类
k <- kmeans(df, centers = 4, nstart = 50)

# 对数据做PCA降维(和fviz_cluster默认逻辑一致)
pca_res <- prcomp(df, scale. = TRUE)
# 提取前两个主成分,合并聚类标签和数据点名称
plot_df <- data.frame(
  PC1 = pca_res$x[,1],
  PC2 = pca_res$x[,2],
  cluster = as.factor(k$cluster),
  name = rownames(df)
)

# 绘制散点图,添加悬停文本映射
p <- ggplot(plot_df, aes(
  x = PC1, 
  y = PC2, 
  color = cluster,
  text = paste("数据点名称:", name, "\n聚类编号:", cluster)
)) +
  geom_point(size = 3) +
  scale_color_brewer(palette = "Set2") +
  theme_minimal() +
  labs(x = "主成分1", y = "主成分2", color = "聚类")

# 转换为交互式图,指定悬停显示text内容
ggplotly(p, tooltip = "text")

方法二:修改fviz_cluster的输出对象

如果想保留fviz_cluster的默认处理逻辑,可直接修改其返回的ggplot数据:

# 加载包
library(factoextra)
library(plotly)
library(stats)
library(ggplot2)

# 读取数据
df <- read.csv("你的测试数据.csv", row.names = 1)

# 执行Kmeans聚类并生成基础图
k <- kmeans(df, centers = 4, nstart = 50)
p_base <- fviz_cluster(k, data=df, palette = "Set2", ggtheme = theme_minimal(), geom = "point")

# 提取fviz_cluster使用的绘图数据,添加数据点名称
plot_df <- p_base$data
plot_df$name <- rownames(df)

# 重新构建ggplot对象,添加悬停文本映射
p <- ggplot(plot_df, aes(
  x = x, 
  y = y, 
  color = cluster,
  text = paste("数据点名称:", name, "\n聚类编号:", cluster)
)) +
  geom_point(size = 3) +
  scale_color_brewer(palette = "Set2") +
  theme_minimal() +
  labs(x = p_base$labels$x, y = p_base$labels$y, color = p_base$labels$color)

# 转换为交互式图
ggplotly(p, tooltip = "text")

两种方法都能实现:仅在鼠标悬停时显示数据点名称,平时只显示聚类散点,同时保留聚类颜色区分的效果。

内容的提问来源于stack exchange,提问作者walter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 04:50:35