关于R语言rep函数及var1_mean/var2_mean箱线图的技术问询
箱线图代码中的rep函数解释及疑问解答
首先是你的样本数据代码:
x1 <- c(1, 2, 3, 4, 5) x2 <- c(6, 7, 4, 5, 7) x3 <- c(4, 5, 3, 7, 1) x4 <- c(3, 5, 6, 4, 2) x5 <- c(1, 3, 4, 4, 2) x6 <- c(4, 5, 4, 3, 5) df <- data.frame(x1 = x1, x2 = x2, x3 = x3, x4 = x4, x5 = x5, x6 = x6) df <- df %>% rowwise() %>% mutate( var1_mean = mean(c(x1, x2, x3)), var2_mean = mean(c(x4, x5, x6)) )
关于rep函数的解释
你提到的c(rep("var1_mean", nrow(df)), rep("var2_mean", nrow(df)))作用如下:
rep(string, times)是R中重复生成元素的函数,第一个参数是要重复的内容,第二个参数是重复次数rep("var1_mean", nrow(df))会生成一个长度等于df行数的向量,每个元素都是"var1_mean"(这里df有5行,所以生成5个"var1_mean")rep("var2_mean", nrow(df))同理生成5个"var2_mean"- 两者拼接后得到一个长度为10的向量,用来给
plot_df的Variable列赋值,让df$var1_mean的每一个数值都对应分组标签"var1_mean",df$var2_mean的每一个数值都对应"var2_mean"——这是ggplot绘制分组图表必须的长格式数据结构。
关于是否计算均值和标准差的疑问
这段代码本身不会主动计算var1_mean和var2_mean的均值、标准差,它只是把原数据中已经存在的var1_mean、var2_mean的所有数值,整理成ggplot能识别的长格式。
箱线图展示的均值、标准差、四分位数等统计量,是由ggplot2的geom_boxplot()自动计算的——它会根据Variable列的分组,对每组的Value列数值统计计算后绘制箱线。
完整绘图代码及示例图
绘图代码:
plot_df <- data.frame( Variable = c(rep("var1_mean", nrow(df)), rep("var2_mean", nrow(df))), Value = c(df$var1_mean, df$var2_mean) ) ggplot(plot_df, aes(x = Variable, y = Value, fill = Variable)) + geom_boxplot() + labs(x = "", y = "Mean Value") + ggtitle("Box Plot of var1_mean and var2_mean") + theme_minimal()
绘制出的箱线图示例:
内容的提问来源于stack exchange,提问作者Espejito
相关产品推荐
相关产品推荐

