You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R新手技术问询:如何用log2转换与stat_summary绘制双因子箱线图

Hey there! Great job getting up and running with ggplot2 and two-factor boxplots—let's walk through how to integrate log2 transformation and stat_summary to make your plot even more informative. Here's a tailored breakdown for your code:

1. Handling Log2 Transformation

Log2 scaling is perfect for data with a wide range of values, but we need to watch out for zero values (since log2(0) is undefined). There are two straightforward approaches:

Option A: Pre-transform your data (add a tiny offset for zeros)

If you want to modify the data directly, add a small value like 1e-6 to avoid infinite values:

# First, create a proper data frame from your vectors
df <- data.frame(
  Systems = Systems,
  treatment = treatment,
  Performance = Performance
)

# Add log2-transformed column (with offset)
df$log2_Perf <- log2(df$Performance + 1e-6)

Option B: Transform on the fly with scale_y_continuous

This keeps your original data intact and handles infinite values automatically:

scale_y_continuous(
  trans = "log2",
  breaks = scales::log_breaks(n = 5), # Nice log-scaled ticks
  labels = scales::label_number(),
  oob = scales::squish_infinite # Squashes any -Inf values to the plot's lower limit
)

2. Adding Summary Stats with stat_summary

stat_summary lets you overlay custom statistics (like means, medians, or confidence intervals) directly on your boxplots. This is great for highlighting central tendencies beyond the boxplot's median line.

For example, let's add red mean points with standard error bars, aligned perfectly with your grouped boxplots:

stat_summary(
  fun = mean, 
  geom = "point", 
  color = "darkred", 
  size = 2,
  position = position_dodge(width = 0.75) # Matches boxplot's dodge width
) +
stat_summary(
  fun.data = mean_se, # Calculates mean + standard error
  geom = "errorbar", 
  color = "darkred", 
  width = 0.2,
  position = position_dodge(width = 0.75)
)

Full Working Example

Putting it all together with your data (I filled in a bit of dummy data to match lengths):

library(ggplot2)

# Your original data
Systems <- c(rep("A", 23), rep("B", 328), rep("C", 85), rep("D", 25))
treatment <- rep(c("MEAN\n0.035 0.252 0.005 0.032", "MEAN\n0.030 0.213 0.008 0.033"), each = 461)
Performance <- c(
  0.041817315, 0.012366105, 0.008223291, 0.101194094, 0.032095414,
  0.022424856, 0.004272651, 0.02568757, 0.012011461, 0.032519093,
  # Dummy data to match vector lengths
  rnorm(461*2 - 10, mean = 0.03, sd = 0.02)
)

# Create data frame
df <- data.frame(Systems, treatment, Performance)

# Plot with log2 transformation (on-the-fly) and summary stats
ggplot(df, aes(x = Systems, y = Performance, fill = treatment)) +
  geom_boxplot(width = 0.7) +
  # Add mean points and error bars
  stat_summary(fun = mean, geom = "point", color = "darkred", size = 2, position = position_dodge(0.7)) +
  stat_summary(fun.data = mean_se, geom = "errorbar", color = "darkred", width = 0.2, position = position_dodge(0.7)) +
  # Log2 scale for y-axis
  scale_y_continuous(
    trans = "log2",
    breaks = scales::log_breaks(n = 5),
    labels = scales::label_number(),
    oob = scales::squish_infinite
  ) +
  # Clean up labels
  labs(
    title = "Performance by System and Treatment (Log2 Scale)",
    x = "System",
    y = "Performance (Log2 Scale)",
    fill = "Treatment"
  ) +
  theme_minimal()

Quick Tips

  • If your Performance data has no zeros, you can skip the 1e-6 offset or squish_infinite argument.
  • Swap fun = mean with fun = median if you want to highlight medians instead.
  • Use fun.data = mean_cl_normal to add 95% confidence intervals instead of standard error bars.

内容的提问来源于stack exchange,提问作者Abdl Razzaq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:34:40