You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在散点图中更清晰展示置信区间?R语言plotly实现咨询

置信区间可视化优化方案

问题说明

你拥有如下数据集:

> dput(dt)
structure(list(Odds.Ratio = c(0.34, 0.85, 0.38, 1.34, 0.98, 0.55, 
0.34, 0.25), Lower.Bound.CI = c(0.12, 0.34, 0.33, 0.8, 0.67, 
0.34, 0.22, 0.13), Upper.Bound.CI = c(0.66, 0.98, 0.67, 1.55, 
1.42, 0.77, 0.5, 0.43), Cluster = c(1L, 1L, 1L, 1L, 2L, 2L, 2L, 
2L)), class = "data.frame", row.names = c(NA, -8L))

你希望绘制各Cluster的Odds.Ratio并直观展示置信区间宽度,当前用点大小映射置信区间宽度的方式在数据量较大时不够清晰,以下是更优的可视化方案:


方案1:误差棒(Error Bars)

直接给每个点添加垂直误差棒,明确展示每个Odds.Ratio的置信区间范围,这是展示单个点置信区间最直观的方式。

library(plotly)
dt$Cluster <- as.factor(dt$Cluster)

fig <- plot_ly(
  dt,
  x = ~Cluster,
  y = ~Odds.Ratio,
  type = 'scatter',
  mode = 'markers',
  color = ~Cluster,
  # 添加置信区间误差棒
  error_y = list(
    type = 'data',
    symmetric = FALSE,
    array = ~Upper.Bound.CI - Odds.Ratio,  # 上偏差
    arrayminus = ~Odds.Ratio - Lower.Bound.CI  # 下偏差
  )
) 

fig %>%
  layout(
    xaxis = list(title = 'Cluster'), 
    yaxis = list(title = 'Odds Ratio'), 
    legend = list(title=list(text='<b> Cluster </b>'))
  )

优势:每个点的置信区间范围一目了然,不会因为数据量增大而混淆,能精准对应到单个样本的区间。


方案2:分组误差带(展示Cluster整体置信区间)

如果需要聚焦每个Cluster的整体统计特征(比如均值的置信区间),可以用分组的误差带搭配散点,同时保留单个样本的信息:

library(plotly)
library(dplyr)

# 计算每个Cluster的均值及整体置信区间(示例用均值的CI,可根据需求替换)
cluster_summary <- dt %>%
  group_by(Cluster) %>%
  summarise(
    mean_OR = mean(Odds.Ratio),
    lower_mean_CI = mean(Lower.Bound.CI),
    upper_mean_CI = mean(Upper.Bound.CI)
  ) %>%
  mutate(Cluster = as.factor(Cluster))

dt$Cluster <- as.factor(dt$Cluster)

fig <- plot_ly() %>%
  # 添加单个样本散点(抖动避免重叠)
  add_trace(
    data = dt,
    x = ~Cluster,
    y = ~Odds.Ratio,
    type = 'scatter',
    mode = 'markers',
    color = ~Cluster,
    opacity = 0.6,
    name = '单个样本'
  ) %>%
  # 添加Cluster均值及误差带
  add_trace(
    data = cluster_summary,
    x = ~Cluster,
    y = ~mean_OR,
    type = 'scatter',
    mode = 'markers+lines',
    marker = list(size = 12, symbol = 'diamond'),
    error_y = list(
      type = 'data',
      symmetric = FALSE,
      array = ~upper_mean_CI - mean_OR,
      arrayminus = ~mean_OR - lower_mean_CI,
      color = 'black'
    ),
    name = 'Cluster均值'
  ) %>%
  layout(
    xaxis = list(title = 'Cluster'), 
    yaxis = list(title = 'Odds Ratio'), 
    legend = list(title=list(text='<b> 图例 </b>'))
  )

优势:同时展示单个样本和Cluster整体的置信区间特征,避免散点重叠,适合数据量较大的场景。


方案3:分组箱线图/小提琴图

如果更关注Cluster内Odds.Ratio的分布及离散程度,箱线图或小提琴图可以结合置信区间展示整体分布:

library(plotly)
dt$Cluster <- as.factor(dt$Cluster)

# 箱线图示例
fig <- plot_ly(dt, x = ~Cluster, y = ~Odds.Ratio, color = ~Cluster, type = 'box') %>%
  # 可选:叠加单个点的置信区间误差棒
  add_trace(
    x = ~Cluster,
    y = ~Odds.Ratio,
    type = 'scatter',
    mode = 'markers',
    error_y = list(
      type = 'data',
      symmetric = FALSE,
      array = ~Upper.Bound.CI - Odds.Ratio,
      arrayminus = ~Odds.Ratio - Lower.Bound.CI
    ),
    opacity = 0.7
  ) %>%
  layout(
    xaxis = list(title = 'Cluster'), 
    yaxis = list(title = 'Odds Ratio'), 
    legend = list(title=list(text='<b> Cluster </b>'))
  )

优势:直观展示Cluster内数据的分布特征,结合误差棒后能同时呈现单个样本的置信区间,适合对比不同Cluster的整体差异。


内容的提问来源于stack exchange,提问作者Jamie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 00:30:01