如何调整Seaborn堆叠概率分布图,实现分类类型分组且每组占比达100%
Got it, let's work through this to get the exact plot you want—stacked bars where each TYPE sums to 100%, with TYPE treated as a categorical variable instead of numeric.
The Problem with Your Original Code
Your sns.displot() call isn't hitting the mark because:
stat="probability"calculates probabilities across the entire dataset, not within eachTYPEgroup. So the stacked bars won't sum to 100% per type.- Even with
discrete=1, ifTYPEis stored as a numeric column, Seaborn will still treat it as a quantitative variable (not categorical), which can lead to unintended spacing or plotting behavior.
Step-by-Step Solution
1. Convert TYPE to a Categorical Variable
First, make sure Seaborn recognizes TYPE as distinct categories (1, 2, 3) instead of numbers:
df["TYPE"] = df["TYPE"].astype("category")
2. Precompute Grouped Percentages
To guarantee each TYPE's bars sum to 100%, we need to calculate the percentage of P and F within each TYPE first:
# Calculate raw counts for each TYPE + GROUP combination counts = df.groupby(["TYPE", "GROUP"]).size().reset_index(name="count") # Compute percentage of each GROUP relative to its parent TYPE counts["percentage"] = counts.groupby("TYPE")["count"].transform(lambda x: x / x.sum() * 100)
3. Plot the Stacked Percentage Bars
Now use sns.catplot() to build the stacked bar chart with our precomputed percentages:
import seaborn as sns import matplotlib.pyplot as plt # Create the plot g = sns.catplot( data=counts, x="TYPE", y="percentage", hue="GROUP", kind="bar", estimator=None, # We already calculated exact percentages, no need to aggregate height=5, aspect=1.2 ) # Customize labels and title g.set_axis_labels("Type", "Percentage (%)") g.fig.suptitle("P/F Group Distribution by Type", y=1.02) # Adjust title position to avoid overlap plt.show()
Alternative: No Precomputation Needed
If you want a more concise approach without precomputing counts, you can use sns.histplot() with stat="percent"—this calculates percentages within each bin (each TYPE) instead of the whole dataset:
df["TYPE"] = df["TYPE"].astype("category") sns.histplot( data=df, x="TYPE", hue="GROUP", multiple="stack", stat="percent", shrink=0.8 # Optional: Adjust bar width for better spacing ) plt.ylabel("Percentage (%)") plt.title("P/F Group Distribution by Type") plt.show()
This works because stat="percent" normalizes each TYPE's counts to 100%, whereas your original stat="probability" normalizes across the entire dataset.
内容的提问来源于stack exchange,提问作者Max

