You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Plotly分组箱线图:自定义分位数后如何对齐显示离群点

解决Plotly分组箱线图中手动添加离群点的对齐问题

由于需要用特定方法(type=5)计算分位数,我手动指定了Plotly箱线图的分位数参数,但添加的离群点散点始终落在分组箱线的中间位置,无法和对应类别的箱线对齐。如何让离群点与所属的箱线精准对齐?

附原效果截图:
分组箱线图(上方已添加离群点)

原可复现代码:

set.seed(123) # Set seed for reproducibility

# Create the site_name column with 5 different site names, each with 20 rows
site_name <- rep(paste0("site_", 1:5), each = 40)

# Create the site_type column with 10 'A's and 10 'B's for each site
site_type <- rep(c("A", "B"), each = 20, times = 5)

# Create the value column with random numbers
value <- runif(100, min = 0, max = 200) # Random numbers between 0 and 200

# Combine into a data frame
df <- data.frame(site_name, site_type, value)

# Display the first few rows of the dataset
head(df, 20)

# Group by site_name and site_type, then calculate summary statistics
stats_df <- df %>%
  group_by(site_name, site_type) %>%
  summarise(
    lower_fence = quantile(value, probs = c(0.05), type = 5, na.rm = TRUE),
    q1 = quantile(value, probs = c(0.25), type = 5, na.rm = TRUE),
    median = quantile(value, probs = c(0.5), type = 5, na.rm = TRUE),
    mean = mean(value, na.rm = TRUE),
    q3 = quantile(value, probs = c(0.75), type = 5, na.rm = TRUE),
    upper_fence = quantile(value, probs = c(0.95), type = 5, na.rm = TRUE),
    sd = sd(value, na.rm = TRUE),
    .groups = 'drop'
  )

# Create the grouped box plot
fig <- plot_ly(
  data = stats_df,
  x = ~factor(site_name),
  color = ~factor(site_type),
  colors = c("blue","red"),
  type = "box",
  source = "boxes",
  lowerfence = ~lower_fence,
  q1 = ~q1,
  median = ~median,
  q3 = ~q3,
  upperfence = ~upper_fence,
  showlegend = TRUE
) %>% 
  layout(boxmode = "group")

# Extract outliers
filtered_df<- df %>%
  left_join(stats_df, by = c("site_name", "site_type")) %>%
  filter(value < lower_fence | value > upper_fence)

# Add the outlier points (原代码此处导致对齐问题)
fig <- fig %>%
  add_trace(
    data = filtered_df,
    x = ~factor(site_name),  
    y = ~value,  
    color = ~factor(site_type),
    colors = c("blue","red"),
    type = "scatter",
    mode = "markers",
    marker = list(size = 5, opacity = 0.6),
    showlegend = FALSE,
    inherit = FALSE
  )

# Show the figure
fig

问题原因

Plotly的分组箱线图(boxmode="group")会自动给同一x分类下的不同color组分配偏移的x坐标,而散点直接使用factor(site_name)会落在该x分类的中心位置,导致和对应箱线错位。

解决方法

将散点的x坐标转换为数值型,并根据site_type添加偏移量,匹配箱线的位置。比如给site_type="A"的点减去0.2,site_type="B"的点加上0.2(偏移量可根据分组数量调整,默认分组箱线的间距是0.4左右)。

修改后的完整代码:

set.seed(123) # Set seed for reproducibility

# Create the site_name column with 5 different site names, each with 20 rows
site_name <- rep(paste0("site_", 1:5), each = 40)

# Create the site_type column with 10 'A's and 10 'B's for each site
site_type <- rep(c("A", "B"), each = 20, times = 5)

# Create the value column with random numbers
value <- runif(100, min = 0, max = 200) # Random numbers between 0 and 200

# Combine into a data frame
df <- data.frame(site_name, site_type, value)

# Display the first few rows of the dataset
head(df, 20)

# Group by site_name and site_type, then calculate summary statistics
stats_df <- df %>%
  group_by(site_name, site_type) %>%
  summarise(
    lower_fence = quantile(value, probs = c(0.05), type = 5, na.rm = TRUE),
    q1 = quantile(value, probs = c(0.25), type = 5, na.rm = TRUE),
    median = quantile(value, probs = c(0.5), type = 5, na.rm = TRUE),
    mean = mean(value, na.rm = TRUE),
    q3 = quantile(value, probs = c(0.75), type = 5, na.rm = TRUE),
    upper_fence = quantile(value, probs = c(0.95), type = 5, na.rm = TRUE),
    sd = sd(value, na.rm = TRUE),
    .groups = 'drop'
  )

# Create the grouped box plot
fig <- plot_ly(
  data = stats_df,
  x = ~factor(site_name),
  color = ~factor(site_type),
  colors = c("blue","red"),
  type = "box",
  source = "boxes",
  lowerfence = ~lower_fence,
  q1 = ~q1,
  median = ~median,
  q3 = ~q3,
  upperfence = ~upper_fence,
  showlegend = TRUE
) %>% 
  layout(boxmode = "group")

# Extract outliers
filtered_df<- df %>%
  left_join(stats_df, by = c("site_name", "site_type")) %>%
  filter(value < lower_fence | value > upper_fence)

# Add the outlier points (修正x坐标偏移)
fig <- fig %>%
  add_trace(
    data = filtered_df,
    # 根据site_type调整x坐标偏移,匹配箱线位置
    x = ~ifelse(site_type == "A", as.numeric(factor(site_name)) - 0.2, as.numeric(factor(site_name)) + 0.2),
    y = ~value,  
    color = ~factor(site_type),
    colors = c("blue","red"),
    type = "scatter",
    mode = "markers",
    marker = list(size = 5, opacity = 0.6),
    showlegend = FALSE,
    inherit = FALSE
  )

# Show the figure
fig

说明

  • 偏移量(0.2)可根据分组数量调整:如果同一x分类下有n个分组,每个分组的偏移量为0.4/(n-1)左右(默认分组箱线的总宽度是0.4)。
  • 转换为数值型x坐标后,Plotly会自动保留原x轴的分类标签,无需额外设置。

内容的提问来源于stack exchange,提问作者TheDza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 14:37:31