You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在面板数据中绘制不同ID间的异质性?R代码优化求助

问题描述

我有如下面板数据:

dput(df2007[1:20, 1:3]) 
structure(list(ID = c(1120L, 1120L, 1120L, 1120L, 1111L, 1111L, 
1111L, 1111L, 1123L, 1123L, 1123L, 1123L, 1135L, 1135L, 1135L, 
1119L, 1119L, 1119L, 1119L, 1124L), yr = c(2007L, 2008L, 2010L, 
2011L, 2007L, 2008L, 2010L, 2011L, 2007L, 2008L, 2010L, 2011L, 
2007L, 2008L, 2010L, 2007L, 2008L, 2010L, 2011L, 2007L), cm = c(1.1, 
1.1, 1.4, 1.3, 1.6, 1.6, 1.7, 1.9, 1.5, 1.5, 1.5, 1.6, 0.9, 1, 
1.2, 2.1, 2.2, 3.4, 4.1, 0.8)), row.names = c("22", "23", "24", 
"25", "171", "172", "173", "174", "214", "215", "216", "217", 
"218", "219", "220", "262", "263", "264", "265", "266"), class = "data.frame")

想要绘制每个ID对应一组散点,同时显示该ID的均值水平线的图形(用来展示不同ID间的异质性),参考网上代码写了如下R代码,但生成的图形十分混乱,求优化方法:

df2007 %>%
  group_by(ID) %>%
  summarise(cm_mean = mean(cm)) %>%
  left_join(df2007) %>%
  ggplot(data = ., 
         aes(x = reorder(as.character(ID), ID), y = cm)) +
  geom_point() +
  geom_line(aes(x = ID, y = cm_mean), col = "blue") +
  labs(x = "ID", y = "Diameter")
优化方案

原代码混乱的核心问题:

  • x轴混用字符型与数值型ID,导致坐标轴不匹配、连线错误
  • 冗余的left_join重复了均值数据,增加不必要的计算
  • geom_line未按ID分组,出现跨ID连线

以下是两种可行的优化思路:

思路1:用stat_summary自动计算并添加均值线

无需提前计算均值,直接借助ggplot的stat_summary工具生成均值水平线,代码更简洁:

library(tidyverse)

df2007 %>%
  ggplot(aes(x = factor(ID), y = cm)) +
  # 给散点添加轻微抖动,避免重叠
  geom_point(position = position_jitter(width = 0.2), alpha = 0.7) +
  # 绘制每个ID的均值水平线
  stat_summary(fun = mean, geom = "hline", aes(yintercept = after_stat(y)), 
               color = "blue", linewidth = 1) +
  labs(x = "ID", y = "Diameter") +
  theme_bw()

如果想让均值标记更直观,可把geom = "hline"换成geom = "crossbar",生成带横线的均值标记。

思路2:提前计算均值,用geom_segment绘制横线

若需要手动控制均值线的长度、样式,可先单独计算均值,再用geom_segment绘制水平线:

library(tidyverse)

# 预计算每个ID的均值
mean_df <- df2007 %>%
  group_by(ID) %>%
  summarise(cm_mean = mean(cm))

df2007 %>%
  ggplot(aes(x = factor(ID), y = cm)) +
  geom_point(position = position_jitter(width = 0.2), alpha = 0.7) +
  # 绘制覆盖ID柱子宽度的均值线
  geom_segment(data = mean_df, 
               aes(x = as.numeric(factor(ID)) - 0.4, 
                   xend = as.numeric(factor(ID)) + 0.4, 
                   y = cm_mean, yend = cm_mean),
               color = "blue", linewidth = 1) +
  labs(x = "ID", y = "Diameter") +
  theme_bw()

额外优化:按均值排序ID

若想强化异质性展示效果,可让ID按均值从小到大排序:

df2007 %>%
  ggplot(aes(x = reorder(factor(ID), cm, FUN = mean), y = cm)) +
  geom_point(position = position_jitter(width = 0.2), alpha = 0.7) +
  stat_summary(fun = mean, geom = "hline", aes(yintercept = after_stat(y)), 
               color = "blue", linewidth = 1) +
  labs(x = "ID (按均值排序)", y = "Diameter") +
  theme_bw()

内容的提问来源于stack exchange,提问作者Jonathen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 09:25:25