Seaborn displot对比不等规模分组密度时估计结果不一致问题求助
Great question! The issue you're seeing is actually due to a default parameter in Seaborn's KDE plotting functions that differs from how pandas and plotnine handle normalization.
Why this happens
By default, sns.displot() (and its underlying sns.kdeplot()) uses common_norm=True. This means it normalizes all the KDE curves together relative to the total combined sample size of all groups. So groups with smaller sample sizes end up with curves that have a smaller total area (proportional to their share of the overall data), which is why your "x" group (100 samples) looks much "thinner" than the "y" group (1000 samples).
In contrast, pandas' plot.kde() and plotnine's geom_density() default to normalizing each group independently, so every curve's area integrates to 1 regardless of sample size—exactly what you expected.
How to fix it
You just need to add the common_norm=False parameter to your displot() call. This tells Seaborn to normalize each group's KDE curve on its own, making their areas equal:
sns.displot(data=dat, x="values", kind="kde", hue="groups", common_norm=False) plt.show()
What this changes
With common_norm=False, both the "x" and "y" curves will have an area of 1, matching the output from pandas and plotnine. You'll see the two curves shifted (as expected from your normal distributions) but with identical "height" scales, making it easy to compare their shapes without the sample size bias.
Bonus: Direct KDE plot alternative
If you prefer using sns.kdeplot() directly (instead of displot()), the same parameter applies:
sns.kdeplot(data=dat, x="values", hue="groups", common_norm=False) plt.show()
This will give you the same corrected result.
内容的提问来源于stack exchange,提问作者ccccc

