如何在R散点图中创建对应分箱范围的组均值分段线?
问题与解决方案
问题描述
我尝试将数据按wdi_expedu划分为等大小分箱,计算每个分箱的wdi_expedu均值,再在wdi_expedu与left_seats的散点图上叠加对应分箱的水平均值分段线。我已经用以下代码创建了4个分箱并生成了组均值数据框:
j <- 4 data_selected_na <- data_selected_na %>% mutate(bin = as.integer(cut_number(wdi_expedu,j))) bin_means <- data_selected_na %>% group_by(bin) %>% summarise(mean_wdi_expedu = mean(wdi_expedu))
但当前绘制的分段线会横跨整个图表,我希望分段线仅覆盖对应分箱的left_seats数值范围,而且不想手动设置坐标(最终j值会扩展到100),请问该怎么实现?当前绘图代码如下:
ggplot(data_selected_na, aes(x = left_seats, y = wdi_expedu)) + geom_point() + geom_segment(data = bin_means, aes(x = min(data_selected_na$left_seats), xend = max(data_selected_na$left_seats), y = bin_means$mean_wdi_expedu, yend = bin_means$mean_wdi_expedu), color = "red") + labs(title = "Scatterplot with Horizontal Lines for j = 4", x = "Left Seats", y = "Education Spending") + theme_minimal()
解决方法
核心是在计算分箱均值时,同步统计每个分箱对应的left_seats极值,用这些极值来限定分段线的x轴范围,无需手动设置坐标。
1. 修正分箱均值计算逻辑
在summarise步骤中新增两个统计量:每个分箱内left_seats的最小值和最大值:
j <- 4 data_selected_na <- data_selected_na %>% mutate(bin = as.integer(cut_number(wdi_expedu,j))) bin_means <- data_selected_na %>% group_by(bin) %>% summarise( mean_wdi_expedu = mean(wdi_expedu), min_left = min(left_seats), # 记录分箱内left_seats的最小值 max_left = max(left_seats) # 记录分箱内left_seats的最大值 )
2. 调整绘图代码
在geom_segment中直接调用bin_means里的min_left和max_left作为x轴起止点,同时避免在aes中使用全局数据引用(已指定data = bin_means,直接用列名即可):
ggplot(data_selected_na, aes(x = left_seats, y = wdi_expedu)) + geom_point() + geom_segment( data = bin_means, aes(x = min_left, xend = max_left, y = mean_wdi_expedu, yend = mean_wdi_expedu), color = "red", linewidth = 1 # 可选:调整线条粗细,提升可读性 ) + labs( title = "散点图:分箱均值水平线(j = 4)", x = "左翼席位", y = "教育支出" ) + theme_minimal()
效果说明
- 修改后,每条红色水平线仅覆盖对应分箱内
left_seats的实际取值区间,不会横跨整个图表。 - 当j扩展到100时,代码无需手动调整,会自动适配所有分箱的极值计算,完全满足批量分箱的需求。
内容的提问来源于stack exchange,提问作者Kate
相关产品推荐
相关产品推荐

