You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将NADA包cenboxplot()生成的截尾数据箱线图转为ggplot风格或小提琴图?

Great question! Working with censored data in ggplot can feel a bit tricky if you’re used to NADA’s cenboxplot(), but we can absolutely replicate that functionality (and level up with violin plots) using ggplot2 and a few helper packages. Let’s walk through how to do both.

First, let’s set up some sample data matching your structure (where ResultCen = 1 means left-censored values below the detection limit, and 0 is a measured value):

set.seed(123) # 保证结果可重复
sample_df <- data.frame(
  Group = rep(c("Control", "Treatment"), each = 60),
  MeasuredValue = c(
    rnorm(50, mean = 15, sd = 3), rep(10, 10), # Control组:50个实测值,10个截尾值(检出限10)
    rnorm(55, mean = 20, sd = 4), rep(12, 5)  # Treatment组:55个实测值,5个截尾值(检出限12)
  ),
  ResultCen = c(rep(0, 50), rep(1, 10), rep(0, 55), rep(1, 5))
)

1. Creating Censored Boxplots in ggplot2

NADA’s cenboxplot() accounts for censored values when calculating quantiles, so we’ll use NADA’s cenquantile() function to compute accurate boxplot stats, then feed those into ggplot2 for full styling control.

library(ggplot2)
library(NADA)
library(dplyr)

# 按分组计算截尾数据的分位数(最小值、25%分位数、中位数、75%分位数、最大值)
censored_quantiles <- sample_df %>%
  group_by(Group) %>%
  summarise(
    ymin = cenquantile(MeasuredValue, ResultCen, p = 0.0),
    lower = cenquantile(MeasuredValue, ResultCen, p = 0.25),
    middle = cenquantile(MeasuredValue, ResultCen, p = 0.5),
    upper = cenquantile(MeasuredValue, ResultCen, p = 0.75),
    ymax = cenquantile(MeasuredValue, ResultCen, p = 1.0)
  )

# 绘制自定义截尾箱线图
ggplot(sample_df, aes(x = Group, y = MeasuredValue)) +
  # 基于截尾分位数的箱线图
  geom_boxplot(
    data = censored_quantiles,
    aes(ymin = ymin, lower = lower, middle = middle, upper = upper, ymax = ymax),
    stat = "identity",
    fill = "#69b3a2",
    alpha = 0.7,
    width = 0.6
  ) +
  # 添加截尾数据点(用向下箭头标记,清晰区分截尾值)
  geom_point(
    data = filter(sample_df, ResultCen == 1),
    shape = 25, # 向下箭头形状
    fill = "#e74c3c",
    size = 2,
    alpha = 0.8
  ) +
  # 添加实测数据点(可选,增强数据可视化)
  geom_jitter(
    data = filter(sample_df, ResultCen == 0),
    shape = 16,
    size = 1,
    alpha = 0.5,
    width = 0.15
  ) +
  labs(
    title = "Censored Data Boxplot (ggplot2)",
    y = "Measured Value",
    x = "Experimental Group"
  ) +
  theme_minimal() +
  theme(plot.title = element_text(hjust = 0.5))

This gives you a fully stylable boxplot that properly accounts for censored values, just like cenboxplot()—but with all the customization power of ggplot2 (colors, themes, labels, etc.).


2. Building Censored Violin Plots

Violin plots show distribution density, which requires handling censored data carefully (since we don’t know the exact values below the detection limit). We’ll fit a censored distribution to each group, then generate the density curve for the violin plot. We’ll use the fitdistrplus package for this.

library(fitdistrplus)
library(purrr)

# 为每个分组拟合左截尾正态分布
censored_fits <- sample_df %>%
  group_split(Group) %>%
  map(function(group_data) {
    # 构建截尾数据格式:left=值,right=NA表示左截尾
    cens_data <- data.frame(
      left = group_data$MeasuredValue,
      right = ifelse(group_data$ResultCen == 1, NA, group_data$MeasuredValue)
    )
    # 拟合正态分布(可根据数据换用其他分布,比如lognormal)
    fitdistcens(cens_data, dist = "norm")
  })

# 生成小提琴图的密度数据(需要对称的密度值来绘制小提琴形状)
violin_density_data <- map_dfr(1:length(censored_fits), function(i) {
  group_name <- unique(sample_df$Group)[i]
  fit <- censored_fits[[i]]
  
  # 生成值的范围
  value_range <- seq(
    min(sample_df$MeasuredValue) - 2,
    max(sample_df$MeasuredValue) + 2,
    length.out = 100
  )
  
  # 计算拟合分布的密度
  density_vals <- dnorm(
    value_range,
    mean = fit$estimate["mean"],
    sd = fit$estimate["sd"]
  )
  
  # 转换为ggplot可用的格式(正负密度实现小提琴对称)
  data.frame(
    Group = group_name,
    Value = value_range,
    Density = density_vals,
    Density_Neg = -density_vals
  )
})

# 绘制截尾小提琴图
ggplot() +
  # 绘制小提琴的上下两部分
  geom_polygon(
    data = violin_density_data,
    aes(x = Group, y = Value, group = Group, fill = Group),
    stat = "identity",
    alpha = 0.6
  ) +
  geom_polygon(
    data = violin_density_data,
    aes(x = Group, y = Density_Neg, group = Group, fill = Group),
    stat = "identity",
    alpha = 0.6
  ) +
  # 添加原始数据点
  geom_jitter(
    data = sample_df,
    aes(x = Group, y = MeasuredValue),
    shape = 16,
    size = 1,
    alpha = 0.5,
    width = 0.15
  ) +
  # 标记截尾数据点
  geom_point(
    data = filter(sample_df, ResultCen == 1),
    aes(x = Group, y = MeasuredValue),
    shape = 25,
    fill = "#e74c3c",
    size = 2,
    alpha = 0.8
  ) +
  labs(
    title = "Censored Data Violin Plot (ggplot2)",
    y = "Measured Value",
    x = "Experimental Group"
  ) +
  theme_minimal() +
  theme(plot.title = element_text(hjust = 0.5)) +
  scale_fill_brewer(palette = "Set2")

This approach creates a violin plot based on the fitted distribution of your censored data, giving you a clear view of the underlying distribution while highlighting censored values.


内容的提问来源于stack exchange,提问作者kslayerr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:52:18