如何基于给定数据使用ggplot绘制抽样分布条形图?
解决方案:用ggplot绘制实际值与样本值的对比条形图
步骤1:数据预处理
首先给原始数据标记出「实际值」和「样本值」,方便绘图时区分:
library(dplyr) library(ggplot2) # 加载原始数据集 set.seed(1) dat = rbind(data.frame('Category' = 'A', values = c(10, rnorm(10000))), data.frame('Category' = 'B', values = c(20, rt(10000, 10)))) # 标记数据类型:每个分类的第一行是实际值,其余为样本值 dat <- dat %>% group_by(Category) %>% mutate(type = ifelse(row_number() == 1, "Actual", "Sample")) %>% ungroup()
步骤2:两种常见绘图方案
根据你需要的展示效果,提供两种实用实现方式:
方案1:实际值条形 + 样本分布箱线图
这种方式既能突出实际值,又能直观展示样本的整体分布情况:
ggplot(dat, aes(x = Category, y = values)) + # 绘制样本值的箱线图,放在底层 geom_boxplot(data = filter(dat, type == "Sample"), width = 0.4, fill = "#E6E6E6") + # 绘制实际值的条形,用醒目颜色区分 geom_col(data = filter(dat, type == "Actual"), aes(fill = type), width = 0.2) + # 给实际值添加数值标签 geom_text(data = filter(dat, type == "Actual"), aes(label = values), vjust = -0.5, size = 4) + # 设置颜色和主题 scale_fill_manual(values = c("Actual" = "#FF6B6B")) + labs(title = "实际值与样本值分布对比", y = "数值", x = "分类", fill = "数据类型") + theme_minimal()
方案2:实际值与样本均值的双条形对比
如果只需要对比实际值和样本的平均水平,可采用双条形图:
# 计算每个分类的样本均值 sample_summary <- dat %>% filter(type == "Sample") %>% group_by(Category) %>% summarise(value = mean(values)) %>% mutate(type = "样本均值") # 提取实际值数据 actual_data <- dat %>% filter(type == "Actual") %>% select(Category, values, type) %>% rename(value = values) # 合并绘图数据 plot_data <- bind_rows(actual_data, sample_summary) # 绘制双条形图 ggplot(plot_data, aes(x = Category, y = value, fill = type)) + geom_col(position = position_dodge(width = 0.8), width = 0.7) + # 添加数值标签 geom_text(aes(label = round(value, 2)), position = position_dodge(width = 0.8), vjust = -0.5, size = 4) + # 设置颜色 scale_fill_manual(values = c("Actual" = "#FF6B6B", "样本均值" = "#4ECDC4")) + labs(title = "实际值与样本均值对比", y = "数值", x = "分类", fill = "数据类型") + theme_minimal()
内容的提问来源于stack exchange,提问作者Bogaso
相关产品推荐
相关产品推荐

