You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整Seaborn堆叠概率分布图,实现分类类型分组且每组占比达100%

Fixing Your Stacked Percentage Bar Plot with Seaborn

Got it, let's work through this to get the exact plot you want—stacked bars where each TYPE sums to 100%, with TYPE treated as a categorical variable instead of numeric.

The Problem with Your Original Code

Your sns.displot() call isn't hitting the mark because:

  • stat="probability" calculates probabilities across the entire dataset, not within each TYPE group. So the stacked bars won't sum to 100% per type.
  • Even with discrete=1, if TYPE is stored as a numeric column, Seaborn will still treat it as a quantitative variable (not categorical), which can lead to unintended spacing or plotting behavior.

Step-by-Step Solution

1. Convert TYPE to a Categorical Variable

First, make sure Seaborn recognizes TYPE as distinct categories (1, 2, 3) instead of numbers:

df["TYPE"] = df["TYPE"].astype("category")

2. Precompute Grouped Percentages

To guarantee each TYPE's bars sum to 100%, we need to calculate the percentage of P and F within each TYPE first:

# Calculate raw counts for each TYPE + GROUP combination
counts = df.groupby(["TYPE", "GROUP"]).size().reset_index(name="count")

# Compute percentage of each GROUP relative to its parent TYPE
counts["percentage"] = counts.groupby("TYPE")["count"].transform(lambda x: x / x.sum() * 100)

3. Plot the Stacked Percentage Bars

Now use sns.catplot() to build the stacked bar chart with our precomputed percentages:

import seaborn as sns
import matplotlib.pyplot as plt

# Create the plot
g = sns.catplot(
    data=counts,
    x="TYPE",
    y="percentage",
    hue="GROUP",
    kind="bar",
    estimator=None,  # We already calculated exact percentages, no need to aggregate
    height=5,
    aspect=1.2
)

# Customize labels and title
g.set_axis_labels("Type", "Percentage (%)")
g.fig.suptitle("P/F Group Distribution by Type", y=1.02)  # Adjust title position to avoid overlap
plt.show()

Alternative: No Precomputation Needed

If you want a more concise approach without precomputing counts, you can use sns.histplot() with stat="percent"—this calculates percentages within each bin (each TYPE) instead of the whole dataset:

df["TYPE"] = df["TYPE"].astype("category")

sns.histplot(
    data=df,
    x="TYPE",
    hue="GROUP",
    multiple="stack",
    stat="percent",
    shrink=0.8  # Optional: Adjust bar width for better spacing
)
plt.ylabel("Percentage (%)")
plt.title("P/F Group Distribution by Type")
plt.show()

This works because stat="percent" normalizes each TYPE's counts to 100%, whereas your original stat="probability" normalizes across the entire dataset.

内容的提问来源于stack exchange,提问作者Max

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 15:28:15