如何在R语言中对坐标数据进行随机分类?
解决方案
1. 数据随机分组处理
先修正原代码中runif的冗余参数,再添加随机分组逻辑:通过随机生成互斥的分组标识,得到Group1和Group2列(1表示属于该组,-1表示不属于)。
完整代码如下:
# Number of observations n <- 250 # x randomly drawn from a continuous uniform distribution with bounds [0,1] x <- runif(min = 0, max = 1, n = n) # Error term from Normal distribution error <- rnorm(n = n, mean = 0, sd = 2) beta_0 <- 1 beta_1 <- -1 y <- beta_0*x + (beta_1*x - error) library(tibble) # 设置随机种子保证分组结果可重复 set.seed(123) # 生成Group1的分组标识,再通过相反数得到互斥的Group2 group1 <- sample(c(1, -1), n, replace = TRUE) group2 <- -group1 # 组装包含分组列的数据表 df <- tibble(x = x, y = y, Group1 = group1, Group2 = group2) df
2. 分组散点图绘制
修改ggplot代码,通过颜色映射区分两组数据,指定蓝色对应Group1、红色对应Group2:
library(ggplot2) # 生成可读性更强的分组名称列(可选,便于图例展示) df <- df %>% mutate(Group = ifelse(Group1 == 1, "Group1", "Group2")) ggplot(data = df, aes(x = x, y = y, color = Group)) + geom_point(size = 2) + # 手动指定两组颜色 scale_color_manual(values = c("Group1" = "blue", "Group2" = "red")) + labs(title = "y = f(x) (分组散点图)", x = "x", y = "y") + theme_minimal()
替代方案(无需额外分组列)
如果不想新增Group列,可直接基于Group1的取值映射颜色:
ggplot(data = df, aes(x = x, y = y, color = factor(Group1))) + geom_point(size = 2) + scale_color_manual(values = c("-1" = "red", "1" = "blue"), labels = c("Group2", "Group1")) + labs(title = "y = f(x) (分组散点图)", x = "x", y = "y", color = "分组") + theme_minimal()
内容的提问来源于stack exchange,提问作者Prometheus
相关产品推荐
相关产品推荐

