You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于numpy直方图计算分布均值与中位数的高效方法

仅用直方图的分箱(bins)和计数(counts)计算均值与中位数

当然可以,基于直方图的分箱和计数,我们可以通过分箱内数据均匀分布的常规假设,高效计算均值和中位数,具体方法如下:

一、计算均值

核心思路是用每个分箱的中点作为该箱内所有数据的代表值,再以计数为权重做加权平均:

  1. 计算每个分箱的中点:取分箱上下界的平均值
  2. 计算所有分箱中点与对应计数的乘积之和
  3. 用总和除以总样本数得到均值

示例代码(基于numpy):

import numpy as np

# 示例直方图数据
bins = np.array([0, 10, 20, 30, 40])
counts = np.array([5, 12, 8, 3])

# 计算分箱中点
bin_midpoints = (bins[:-1] + bins[1:]) / 2
# 计算加权和与总样本数
weighted_sum = np.sum(bin_midpoints * counts)
total_samples = np.sum(counts)
# 得到均值
mean = weighted_sum / total_samples

print(f"估算均值: {mean:.2f}")

二、计算中位数

中位数是第总样本数/2个位置的数值,需要通过累积计数定位到对应分箱,再做线性插值:

  1. 计算累积计数数组,记录到每个分箱为止的总样本数
  2. 找到包含中位数位置的分箱(第一个累积计数≥中位数位置的分箱)
  3. 在该分箱内,根据剩余需要的样本数占分箱总计数的比例,线性插值计算中位数

示例代码:

# 基于上面的bins和counts继续计算
total_samples = np.sum(counts)
median_position = total_samples / 2
cumulative_counts = np.cumsum(counts)

# 定位中位数所在的分箱索引
median_bin_idx = np.argmax(cumulative_counts >= median_position)
# 获取分箱的上下界
lower_bound = bins[median_bin_idx]
upper_bound = bins[median_bin_idx + 1]
# 获取该分箱之前的累积样本数
prev_cumulative = cumulative_counts[median_bin_idx - 1] if median_bin_idx > 0 else 0

# 线性插值计算中位数
median = lower_bound + ( (median_position - prev_cumulative) / counts[median_bin_idx] ) * (upper_bound - lower_bound)

print(f"估算中位数: {median:.2f}")

注意事项

上述方法的精度依赖于分箱内数据均匀分布的假设,如果原始数据在分箱内分布极度不均,结果会存在偏差,但这是无原始数据时的最优估算方式。

内容的提问来源于stack exchange,提问作者BlackPhoenix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 18:07:08