如何在ggplot2中添加聚类中心点?(KNN教学Shiny应用场景)
在ggplot2中添加聚类中心点的实现方案
嘿,刚好做过类似的KNN教学demo,给你一套清晰的步骤,直接上手就能实现需求:
核心思路
先给iris数据集随机分配聚类标签,然后计算每个聚类的Sepal.Length和Sepal.Width均值(也就是聚类中心),最后在散点图上叠加这些中心点即可。
完整代码示例
我分两种方式写,一种用dplyr(代码更简洁易读,适合教学展示逻辑),一种用base R(不用额外加载包),你可以选自己顺手的:
方式1:用dplyr处理数据(推荐)
# 加载需要的包 library(ggplot2) library(dplyr) # 1. 给iris添加随机聚类标签(0和1) set.seed(123) # 设置随机种子,保证每次运行的聚类结果一致,适合教学演示 iris_clustered <- iris %>% mutate(cluster = sample(c(0, 1), nrow(.), replace = TRUE)) # 2. 计算每个聚类的中心点(取Sepal.Length和Sepal.Width的均值) cluster_centers <- iris_clustered %>% group_by(cluster) %>% summarize( Sepal.Length = mean(Sepal.Length), Sepal.Width = mean(Sepal.Width) ) # 3. 绘制散点图并叠加中心点 ggplot(iris_clustered, aes(x = Sepal.Length, y = Sepal.Width, color = factor(cluster))) + # 绘制原始数据点,加alpha透明度让中心点更突出 geom_point(alpha = 0.7) + # 绘制聚类中心点:设置大尺寸、粗边框,用形状区分不同聚类 geom_point(data = cluster_centers, aes(shape = factor(cluster)), size = 5, stroke = 2) + # 添加标签优化可读性 labs( x = "Sepal Length", y = "Sepal Width", color = "Cluster", shape = "Cluster Center", title = "Iris Data with Random Clusters & Centers" ) + theme_minimal()
方式2:用base R处理数据(无需额外包)
如果不想加载dplyr,用base R的aggregate函数也能快速计算中心点:
library(ggplot2) # 1. 添加随机聚类标签 set.seed(123) iris$cluster <- sample(c(0, 1), nrow(iris), replace = TRUE) # 2. 计算聚类中心点 cluster_centers <- aggregate(cbind(Sepal.Length, Sepal.Width) ~ cluster, data = iris, FUN = mean) # 3. 绘图(和上面逻辑一致) ggplot(iris, aes(x = Sepal.Length, y = Sepal.Width, color = factor(cluster))) + geom_point(alpha = 0.7) + geom_point(data = cluster_centers, aes(shape = factor(cluster)), size = 5, stroke = 2) + labs( x = "Sepal Length", y = "Sepal Width", color = "Cluster", shape = "Cluster Center", title = "Iris Data with Random Clusters & Centers" ) + theme_minimal()
额外优化建议(适合教学展示)
如果想让中心点更直观,可以给每个中心加上文字标签,只需要在绘图代码里加这一行即可:
geom_text(data = cluster_centers, aes(label = paste("Center", cluster)), vjust = -1, size = 4, fontface = "bold")
这样观众一眼就能看出每个中心点对应的聚类编号,教学效果会更好。
内容的提问来源于stack exchange,提问作者Henk Straten
相关产品推荐
相关产品推荐

