Plotly分组箱线图:自定义分位数后如何对齐显示离群点
解决Plotly分组箱线图中手动添加离群点的对齐问题
由于需要用特定方法(type=5)计算分位数,我手动指定了Plotly箱线图的分位数参数,但添加的离群点散点始终落在分组箱线的中间位置,无法和对应类别的箱线对齐。如何让离群点与所属的箱线精准对齐?
附原效果截图:
原可复现代码:
set.seed(123) # Set seed for reproducibility # Create the site_name column with 5 different site names, each with 20 rows site_name <- rep(paste0("site_", 1:5), each = 40) # Create the site_type column with 10 'A's and 10 'B's for each site site_type <- rep(c("A", "B"), each = 20, times = 5) # Create the value column with random numbers value <- runif(100, min = 0, max = 200) # Random numbers between 0 and 200 # Combine into a data frame df <- data.frame(site_name, site_type, value) # Display the first few rows of the dataset head(df, 20) # Group by site_name and site_type, then calculate summary statistics stats_df <- df %>% group_by(site_name, site_type) %>% summarise( lower_fence = quantile(value, probs = c(0.05), type = 5, na.rm = TRUE), q1 = quantile(value, probs = c(0.25), type = 5, na.rm = TRUE), median = quantile(value, probs = c(0.5), type = 5, na.rm = TRUE), mean = mean(value, na.rm = TRUE), q3 = quantile(value, probs = c(0.75), type = 5, na.rm = TRUE), upper_fence = quantile(value, probs = c(0.95), type = 5, na.rm = TRUE), sd = sd(value, na.rm = TRUE), .groups = 'drop' ) # Create the grouped box plot fig <- plot_ly( data = stats_df, x = ~factor(site_name), color = ~factor(site_type), colors = c("blue","red"), type = "box", source = "boxes", lowerfence = ~lower_fence, q1 = ~q1, median = ~median, q3 = ~q3, upperfence = ~upper_fence, showlegend = TRUE ) %>% layout(boxmode = "group") # Extract outliers filtered_df<- df %>% left_join(stats_df, by = c("site_name", "site_type")) %>% filter(value < lower_fence | value > upper_fence) # Add the outlier points (原代码此处导致对齐问题) fig <- fig %>% add_trace( data = filtered_df, x = ~factor(site_name), y = ~value, color = ~factor(site_type), colors = c("blue","red"), type = "scatter", mode = "markers", marker = list(size = 5, opacity = 0.6), showlegend = FALSE, inherit = FALSE ) # Show the figure fig
问题原因
Plotly的分组箱线图(boxmode="group")会自动给同一x分类下的不同color组分配偏移的x坐标,而散点直接使用factor(site_name)会落在该x分类的中心位置,导致和对应箱线错位。
解决方法
将散点的x坐标转换为数值型,并根据site_type添加偏移量,匹配箱线的位置。比如给site_type="A"的点减去0.2,site_type="B"的点加上0.2(偏移量可根据分组数量调整,默认分组箱线的间距是0.4左右)。
修改后的完整代码:
set.seed(123) # Set seed for reproducibility # Create the site_name column with 5 different site names, each with 20 rows site_name <- rep(paste0("site_", 1:5), each = 40) # Create the site_type column with 10 'A's and 10 'B's for each site site_type <- rep(c("A", "B"), each = 20, times = 5) # Create the value column with random numbers value <- runif(100, min = 0, max = 200) # Random numbers between 0 and 200 # Combine into a data frame df <- data.frame(site_name, site_type, value) # Display the first few rows of the dataset head(df, 20) # Group by site_name and site_type, then calculate summary statistics stats_df <- df %>% group_by(site_name, site_type) %>% summarise( lower_fence = quantile(value, probs = c(0.05), type = 5, na.rm = TRUE), q1 = quantile(value, probs = c(0.25), type = 5, na.rm = TRUE), median = quantile(value, probs = c(0.5), type = 5, na.rm = TRUE), mean = mean(value, na.rm = TRUE), q3 = quantile(value, probs = c(0.75), type = 5, na.rm = TRUE), upper_fence = quantile(value, probs = c(0.95), type = 5, na.rm = TRUE), sd = sd(value, na.rm = TRUE), .groups = 'drop' ) # Create the grouped box plot fig <- plot_ly( data = stats_df, x = ~factor(site_name), color = ~factor(site_type), colors = c("blue","red"), type = "box", source = "boxes", lowerfence = ~lower_fence, q1 = ~q1, median = ~median, q3 = ~q3, upperfence = ~upper_fence, showlegend = TRUE ) %>% layout(boxmode = "group") # Extract outliers filtered_df<- df %>% left_join(stats_df, by = c("site_name", "site_type")) %>% filter(value < lower_fence | value > upper_fence) # Add the outlier points (修正x坐标偏移) fig <- fig %>% add_trace( data = filtered_df, # 根据site_type调整x坐标偏移,匹配箱线位置 x = ~ifelse(site_type == "A", as.numeric(factor(site_name)) - 0.2, as.numeric(factor(site_name)) + 0.2), y = ~value, color = ~factor(site_type), colors = c("blue","red"), type = "scatter", mode = "markers", marker = list(size = 5, opacity = 0.6), showlegend = FALSE, inherit = FALSE ) # Show the figure fig
说明
- 偏移量(0.2)可根据分组数量调整:如果同一x分类下有n个分组,每个分组的偏移量为
0.4/(n-1)左右(默认分组箱线的总宽度是0.4)。 - 转换为数值型x坐标后,Plotly会自动保留原x轴的分类标签,无需额外设置。
内容的提问来源于stack exchange,提问作者TheDza
相关产品推荐
相关产品推荐

