You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否用Seaborn的ecdfplot或兼容displot的函数展示变量集中度?

需求场景

我希望绘制图表展示某一变量的集中度,例如我有一维价格数组:

  • 需要展示最贵的前10件商品占总价的10%、前100件占总价的40%这类信息
  • 这类可视化适用于分析数据集中程度的场景,比如少数借款人占据银行大部分风险敞口、少数天数占据某时段大部分降雨量等
当前实现方式

我手动按价格排序、计算累积和并除以总价后绘图,但这种方式不够理想。

期望改进方向

希望使用Seaborn的displot和facetgrids实现多分类下的此类计算,类似参考图效果。

核心问题

能否使用ecdfplot或其他兼容Seaborn displot的函数实现该需求?

现有可行但不理想的代码
import numpy as np
from numpy.random import default_rng
import pandas as pd
import copy

import matplotlib
matplotlib.use('TkAgg', force = True)
import matplotlib.pyplot as plt

import seaborn as sns
import seaborn.objects as so
from matplotlib.ticker import FuncFormatter
sns.set_style("darkgrid")
rng = default_rng()

# I generate random samples from a truncated normal distr
# (I don't want negative values)
n = int(2e3)
n_red = int(n/3)
n_green = n - n_red
df = pd.DataFrame()
df['price']= np.random.randn(n) * 100 + 20
df['colour'] = np.hstack([np.repeat('red',n_red),
                          np.repeat('green', n_green)])
df = copy.deepcopy(df.query('price > 0')).reset_index(drop=True)

num_cols = len(np.unique(df['colour']))
fig1, ax1 = plt.subplots(num_cols)

sub_dfs={}
for my_ax, c in enumerate(np.unique(df['colour'])):
    sub_dfs[c] = copy.deepcopy(df.query('colour == @c'))
    sub_dfs[c] = sub_dfs[c].sort_values(by='price', ascending=False).reset_index()
    sub_dfs[c]['cum %'] = np.cumsum(sub_dfs[c]['price']) / sub_dfs[c]['price'].sum()

    sns.lineplot(sub_dfs[c]['cum %'], ax = ax1[my_ax])
    ax1[my_ax].set_title(c + ' - price concentration')
    ax1[my_ax].set_xlabel('# of items')
    ax1[my_ax].set_ylabel('% of total price')
尝试过但无效的代码

我尝试了displot和ecdf相关代码:

fig2 = sns.displot(kind='ecdf', data = df, y='price', col='colour', col_wrap =2, weights ='price',
                   facet_kws=dict(sharey=False))

fig3 = sns.displot(kind='ecdf', data = df, x='price', col='colour', col_wrap =2, weights='price',
                   facet_kws=dict(sharey=False))
补充:按建议尝试仍未解决的问题

@mwaskom,感谢您的解答,但我仍未得到预期结果:

运行以下代码:

fig5 = sns.displot(kind='ecdf', data=df, x=df.index, col='colour', col_wrap =2, weights='price',
                   facet_kws=dict(sharey=False, sharex=False))
  • 得到的是两条直线,而我需要的是凸线(直线代表价格均匀分布,凸线代表前N%商品占比超N%,符合我的数据情况)
  • 示例中一类约400条数据、一类约800条,但x轴使用整个DataFrame的索引,导致第二个子图x轴范围为400-1200而非1-800

内容的提问来源于stack exchange,提问作者Pythonista anonymous

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 17:15:04