You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

带散点的箱线图:类别结果相近时避免重叠的方法

解决箱线图重叠且散点错位的问题

我正在绘制包含不同类别的箱线图,并叠加散点。当不同类别的结果非常相近时(如图中B面板的TSS指标),箱线图会出现重叠。尝试调整箱线宽度会导致散点与箱线错位,效果不理想,求能保持散点与箱线对齐同时避免重叠的方法。

现有代码:

map_levels = {
    0.5: "High",
    0.2: "Medium",
    0.1: "Low",
    0.01: "Extremely low"
}

simulated_RF["Prevalence_level"] = simulated_RF["Prevalence"].map(map_levels)

order_levels = ["High", "Medium", "Low", "Extremely low"]
simulated_RF["Prevalence_level"] = pd.Categorical(
    simulated_RF["Prevalence_level"], categories=order_levels, ordered=True
)


metrics_to_plot = ["AUC", "TSS", "BrierScore", "LogLoss"]

palette_custom  = ["cornflowerblue", "orange"]

fig, axes = plt.subplots(2, 2, figsize=(14, 10))
axes = axes.flatten()

panel_labels = ["A)", "B)", "C)", "D)"]

for i, metric in enumerate(metrics_to_plot):
    
    ax = axes[i]
    df_m = simulated_RF[simulated_RF["Metric"] == metric]
    
    
    sns.boxplot(
        data=df_m,
        x="Prevalence_level",
        y="Value",
        hue="Type",
        dodge=True,
        ax=ax,
        palette = palette_custom 
    )
    
    
    sns.stripplot(
        data=df_m,
        x="Prevalence_level",
        y="Value",
        hue="Type",
        dodge=True,
        palette=["black", "black"],
        size=4,
        jitter=True,
        alpha=0.5,
        ax=ax
    )
    
   
    ax.set_title(metric, fontsize=16)
    ax.set_xlabel("Prevalence level", fontsize=13)
    ax.set_ylabel(metric, fontsize=13)
    ax.set_xticks(range(len(order_levels)))
    ax.set_xticklabels(order_levels, fontsize=11)
    ax.tick_params(axis="y", labelsize=11)
    
    ax.text(-0.1, 1.05, panel_labels[i],
            transform=ax.transAxes, fontsize=16, fontweight="bold")
    
    handles, labels = ax.get_legend_handles_labels()
    if i == 0:
        legend_handles = handles[:2]
        legend_labels = labels[:2]
    ax.get_legend().remove()

fig.legend(
    legend_handles, legend_labels,
    loc="lower center", ncol=2,
    fontsize=13, title_fontsize=13
)

fig.tight_layout(rect=[0, 0.05, 1, 0.95])
plt.show() 

解决方案

方法1:自定义偏移量同步箱线与散点

将boxplot和stripplot的dodge参数从True改为统一的数值(比如0.3),手动控制分组间的偏移距离,既拉开箱线避免重叠,又保证散点与箱线严格对齐。数值可根据图表宽度调整,越大偏移越明显。

修改后的核心代码片段:

sns.boxplot(
    data=df_m,
    x="Prevalence_level",
    y="Value",
    hue="Type",
    dodge=0.3,  # 替换True为具体数值
    ax=ax,
    palette=palette_custom
)

sns.stripplot(
    data=df_m,
    x="Prevalence_level",
    y="Value",
    hue="Type",
    dodge=0.3,  # 与boxplot的dodge值保持一致
    palette=["black", "black"],
    size=4,
    jitter=True,
    alpha=0.5,
    ax=ax
)

方法2:使用catplot统一管理布局

用seaborn的catplot可以自动同步箱线和散点的偏移逻辑,避免手动调整的错位问题,代码更简洁,布局控制更统一:

map_levels = {
    0.5: "High",
    0.2: "Medium",
    0.1: "Low",
    0.01: "Extremely low"
}

simulated_RF["Prevalence_level"] = simulated_RF["Prevalence"].map(map_levels)

order_levels = ["High", "Medium", "Low", "Extremely low"]
simulated_RF["Prevalence_level"] = pd.Categorical(
    simulated_RF["Prevalence_level"], categories=order_levels, ordered=True
)

metrics_to_plot = ["AUC", "TSS", "BrierScore", "LogLoss"]
palette_custom = ["cornflowerblue", "orange"]

# 使用catplot创建分面图
g = sns.catplot(
    data=simulated_RF[simulated_RF["Metric"].isin(metrics_to_plot)],
    x="Prevalence_level",
    y="Value",
    hue="Type",
    col="Metric",
    col_wrap=2,
    kind="box",
    dodge=0.3,
    palette=palette_custom,
    height=5,
    aspect=1.4
)

# 在每个子图上叠加散点
for ax in g.axes.flat:
    metric_name = ax.get_title().split("=")[1].strip()
    sns.stripplot(
        data=simulated_RF[simulated_RF["Metric"] == metric_name],
        x="Prevalence_level",
        y="Value",
        hue="Type",
        dodge=0.3,
        palette=["black", "black"],
        size=4,
        jitter=True,
        alpha=0.5,
        ax=ax
    )
    # 移除子图内的图例,保留全局图例
    ax.get_legend().remove()
    # 添加面板标签
    idx = metrics_to_plot.index(metric_name)
    ax.text(-0.1, 1.05, ["A)", "B)", "C)", "D)"][idx],
            transform=ax.transAxes, fontsize=16, fontweight="bold")

# 设置全局图例
g.add_legend(title="Type", fontsize=13, title_fontsize=13)
# 调整布局
g.tight_layout(rect=[0, 0.05, 1, 0.95])
plt.show()

方法3:缩小箱线宽度+同步散点偏移

如果倾向于调整箱线宽度而非偏移,需同时设置boxplot的width参数和两者的dodge数值,保证散点与箱线位置匹配:

修改后的核心代码片段:

sns.boxplot(
    data=df_m,
    x="Prevalence_level",
    y="Value",
    hue="Type",
    dodge=0.4,  # 与width值对应
    width=0.4,  # 缩小箱线宽度
    ax=ax,
    palette=palette_custom
)

sns.stripplot(
    data=df_m,
    x="Prevalence_level",
    y="Value",
    hue="Type",
    dodge=0.4,  # 与boxplot的dodge值一致
    palette=["black", "black"],
    size=4,
    jitter=True,
    alpha=0.5,
    ax=ax
)

内容的提问来源于stack exchange,提问作者Lola Riesgo Torres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 01:50:57