基于R ggplot绘制分试验得分区间计数直方图的问题求助
问题描述
我有一个包含三列的数据集:参与者编号、试验编号、参与者在该试验中的得分。数据集涉及100名参与者与160次试验,我希望绘制一幅直方图,统计每个试验中得分处于0-2、2-4、4-6……以及大于16区间的参与者数量。
我尝试了如下代码:
max(scores.data$value) scores.data %>% ggplot(aes(score, fill = trial)) + geom_bar(color = NA, position = position_dodge()) + scale_y_continuous(expand = c(0,0), limits = c(0, 100), breaks = seq(0, 100, by = 5)) + scale_x_binned(limits = c(0,45), breaks = c(2,4,6,8,10,12,14,16,45), labels = c("2","4","6","8","10","12","14","16","45")) + labs(x = "time bins per trial", y = "count" ) + facet_grid(~trial, scales = "free_x", switch = "x") + theme_classic() + theme(axis.text.x = element_blank(), axis.ticks.x = element_blank(), legend.position = "none", panel.grid.major.y = element_line(color = "lightgray", size = 0.25), panel.spacing = unit(0, "points"), strip.background = element_blank(), strip.text = element_blank())
但生成的图不符合预期:
以下是可复现数据:
structure(list(participant = c("10000", "10000", "10000", "10000", "10000", "10000", "10000", "10000", "10000", "10000"), trial = c("1", "2", "3", "4", "5", "6", "7", "8", "9", "10"), score = c(5.409, 4.079, 4.355, 4.245, 3.43, 3.685, 4.808, 3.256, 7.038, 3.714)), row.names = c(NA, -10L), class = c("tbl_df", "tbl", "data.frame"))
解决方案
原代码的核心问题:geom_bar搭配position_dodge会错误错开同区间的柱子,且facet_grid的设置隐藏了试验标签,未正确展示每个试验的得分区间分布。以下是两种可行的修正方案:
方法1:手动分箱统计后绘图(可控性强)
先对数据按试验和得分区间做统计,再用柱状图展示:
library(tidyverse) # 定义得分区间与对应标签 score_bins <- c(0, 2, 4, 6, 8, 10, 12, 14, 16, Inf) bin_labels <- c("0-2", "2-4", "4-6", "6-8", "8-10", "10-12", "12-14", "14-16", ">16") # 分箱并统计每个试验各区间的参与者数量 scores_summary <- scores.data %>% mutate(score_bin = cut(score, breaks = score_bins, labels = bin_labels, right = FALSE)) %>% group_by(trial, score_bin) %>% summarise(count = n(), .groups = "drop") %>% # 补全无数据的区间,保证所有子图x轴一致 complete(trial, score_bin, fill = list(count = 0)) # 绘图 ggplot(scores_summary, aes(x = score_bin, y = count, fill = score_bin)) + geom_col(color = "black", width = 0.8) + # 按试验分面,可根据试验总数调整nrow/ncol facet_wrap(~trial, nrow = 10) + scale_y_continuous(expand = c(0, 0), limits = c(0, 100), breaks = seq(0, 100, 5)) + labs(x = "得分区间", y = "参与者数量", title = "各试验得分区间参与者分布") + theme_classic() + theme( legend.position = "none", panel.grid.major.y = element_line(color = "lightgray", size = 0.25), axis.text.x = element_text(angle = 45, hjust = 1) )
方法2:直接用直方图分面(简洁高效)
无需手动统计,直接通过geom_histogram指定分箱规则后分面:
ggplot(scores.data, aes(x = score)) + geom_histogram( breaks = score_bins, color = "black", fill = "steelblue", closed = "left" # 左闭右开区间,匹配0-2包含0、不包含2的需求 ) + facet_wrap(~trial, nrow = 10) + scale_y_continuous(expand = c(0, 0), limits = c(0, 100), breaks = seq(0, 100, 5)) + scale_x_continuous(breaks = score_bins, labels = bin_labels) + labs(x = "得分区间", y = "参与者数量", title = "各试验得分区间参与者分布") + theme_classic() + theme( panel.grid.major.y = element_line(color = "lightgray", size = 0.25), axis.text.x = element_text(angle = 45, hjust = 1) )
关键调整说明
- 明确区间划分:用
cut或breaks指定目标区间,right=FALSE/closed="left"确保区间逻辑符合预期。 - 合理分面展示:用
facet_wrap替代facet_grid,保留试验标签并灵活控制布局。 - 补全缺失区间:手动统计时用
complete补全无数据的区间,避免子图柱子缺失、x轴不一致。 - 移除冗余设置:原代码中
fill=trial和position_dodge完全多余,每个分面已对应单个试验,无需额外区分。
内容的提问来源于stack exchange,提问作者AnneZ
相关产品推荐
相关产品推荐

