You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于直方图分箱输入计算百分位数的实现方法咨询

基于直方图分箱数据的百分位数计算实现

pandas.qcut是基于全量原始样本进行分位划分的接口,你当前只有分箱汇总后的直方图统计结果,没有原始样本数据,因此无法直接用该接口得到正确结果。

实现逻辑

  • 先计算所有分箱的总样本数,依次计算每个分箱的累计样本占比
  • 针对每个目标百分位,先定位其所属的分箱区间
  • 在对应分箱内做线性插值,计算得到百分位对应的数值,边界值(0分位、100分位)直接取分箱的左右端点即可

完整Python实现

def percentile_from_hist(hist, edges, percentiles):
    total_sample = sum(hist)
    # 计算累计样本占比
    cum_pct = []
    current_cum = 0
    for count in hist:
        current_cum += count
        cum_pct.append(current_cum / total_sample)
    
    result = []
    for p in percentiles:
        p_normal = p / 100
        # 0分位直接返回左边界
        if p_normal <= 0:
            result.append(edges[0])
            continue
        # 100分位直接返回右边界
        if p_normal >= 1:
            result.append(edges[-1])
            continue
        # 定位百分位所属分箱
        for bin_idx in range(len(cum_pct)):
            if cum_pct[bin_idx] >= p_normal:
                prev_cum_pct = cum_pct[bin_idx -1] if bin_idx > 0 else 0
                bin_left = edges[bin_idx]
                bin_right = edges[bin_idx + 1]
                bin_total_pct = cum_pct[bin_idx] - prev_cum_pct
                # 线性插值计算结果
                val = bin_left + (p_normal - prev_cum_pct) / bin_total_pct * (bin_right - bin_left)
                result.append(val)
                break
    return result

# 代入你的参数测试
hist = [10, 15, 4]
edges = [0.5, 6, 12, 25]
target_percentiles = [0, 25, 50, 75, 100]
print(percentile_from_hist(hist, edges, target_percentiles))

运行上述代码输出为[0.5, 4.4875, 7.8, 10.7, 25],和你预期的结果完全一致。

内容的提问来源于stack exchange,提问作者Decile

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 06:09:03