You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中numpy、statistics分位数与自定义函数结果差异排查

自定义分位数函数与标准库结果差异的原因

我编写了一个用于计算分位数/十分位数/百分位数的自定义函数:

from math import floor, ceil
def mean(x: list)->float:
    """列表的算术平均值(均值)
    将所有元素求和后除以元素个数"""
    if len(x) < 1:
        raise TypeError("mean need at least one element")
    else:
        sum_of_elements =sum(x)
        mean_value = sum_of_elements/len(x)
        return mean_value

def percentiles(percentile, x:list)->float:
    '''百分位数是指某一百分比的样本值所低于的数值。
    percentile=25、50、75、100对应第1、2、3、4四分位数;
    percentile=10、20…100对应第1至第10十分位数。'''
    x = sorted(x)
    n = len(x)
    if percentile <1 or percentile >100:
        raise TypeError("percentile should be in range 1-100")
    if n<0:
        raise TypeError("list cannot be empty")
    location = percentile*(n+1)/100
    if int(location) == location:
        return x[int(location)-1]
    else:
        down = floor(location)
        up = ceil(location)
        if down == 0:
            return x[up]
        else:
            even = mean([x[down-1], x[up-1]])
            return even

将该函数与np.quantile和statistics.quantiles对比时,得到了不同结果。例如针对列表[1,2,3,4,5,6,7,8,10,11,12,15,16,21]:

  • 自定义函数计算的Q1为3.5;
  • statistics库的Q1为3.75(method='exclusive')或4.25(method='inclusive');
  • numpy库的Q1为4.25。

我的percentiles()函数存在什么问题?我手动计算时结果始终与自定义函数一致。


问题核心:分位数计算有多种算法标准

你的自定义函数没出错,只是和NumPy、statistics库采用了不同的分位数计算规则——分位数的计算从来没有唯一的“标准答案”,不同工具会选用不同的主流算法,这才导致结果差异。

1. 你的函数的计算逻辑

拿你给出的列表(样本量n=14)计算Q1(25分位数):

  • 你用的位置公式:location = 25*(14+1)/100 = 3.75
  • 取第3位和第4位的元素(Python列表索引为2和3):3和4,取平均得到3.5,和你手动计算的结果一致。

2. NumPy默认的算法

NumPy默认采用“基于n-1的线性插值”规则,位置公式为:location = (25/100)*(14-1)
计算得3.25,对应索引3的元素4和索引4的元素5,按权重插值:4 + 0.25*(5-4) = 4.25,这就是你看到的NumPy结果。

3. statistics库的两种算法

  • method='exclusive':采用排除中位数的小样本计算逻辑,位置计算后对应3和4插值,得到3.75
  • method='inclusive':采用包含中位数的逻辑,位置计算后对应4和5插值,得到4.25

对齐标准库的修改方案

如果想让自定义函数和NumPy结果对齐,只需修改位置计算公式:

def percentiles_np_style(percentile, x:list)->float:
    x = sorted(x)
    n = len(x)
    if percentile <1 or percentile >100:
        raise TypeError("percentile should be in range 1-100")
    if n <1:
        raise TypeError("list cannot be empty")
    p = percentile / 100
    location = p*(n-1)
    if location.is_integer():
        return x[int(location)]
    else:
        down_idx = int(floor(location))
        up_idx = int(ceil(location))
        weight = location - down_idx
        return x[down_idx] + weight*(x[up_idx]-x[down_idx])

测试这个版本,输入你的列表,Q1会得到4.25,和NumPy结果一致。


内容的提问来源于stack exchange,提问作者Karolina Jonczyk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 10:16:20