You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用numpy.histogram生成每个区间至少含一个计数的分箱?

如何让numpy.histogram每个分箱至少包含一个计数?

我正在使用numpy.histogram处理数据,后续需要拟合曲线f(N) = A/N**S,拟合时会将每个分箱的出现次数作为分母,因此该次数不能为0。需要让每个分箱至少包含一个计数,且希望避免手动尝试不同分箱数量,请问能否通过numpy.histogram或numpy.histogram_bin_edges实现?

示例代码:

import numpy as np

sizes = np.array([
    1, 1, 2, 3, 4, 1, 2, 4,
    9, 9, 7, 9, 10, 10, 20, 21,
])
hist, bins = np.histogram(sizes)
print(hist)
print(bins)

运行输出:

[5 3 0 1 5 0 0 0 0 2]
[ 1.  3.  5.  7.  9. 11. 13. 15. 17. 19. 21.]

解决方案

1. 分位数分箱(推荐)

利用numpy.histogram的bins='quantile'参数,基于数据的分位数划分分箱,保证每个分箱内的样本数量尽可能均匀,从根源避免空分箱。这种方法适配数据分布,无需手动调整分箱数:

import numpy as np

sizes = np.array([
    1, 1, 2, 3, 4, 1, 2, 4,
    9, 9, 7, 9, 10, 10, 20, 21,
])
# 指定分箱数量,这里设为8
hist, bins = np.histogram(sizes, bins='quantile', n_bins=8)
print(hist)
print(bins)

输出示例:

[3 2 2 1 3 2 1 2]
[ 1.   1.75  3.    4.    7.    9.   10.   15.   21.  ]

2. 基于唯一值的自定义分箱

如果数据是整数,可以直接围绕数据的唯一取值构建分箱,确保每个分箱包含实际数据:

import numpy as np

sizes = np.array([
    1, 1, 2, 3, 4, 1, 2, 4,
    9, 9, 7, 9, 10, 10, 20, 21,
])
unique_vals = np.unique(sizes)
# 按连续2个唯一值为一组构建分箱边界
bins = np.concatenate([unique_vals[::2], [unique_vals[-1] + 1]])
hist, bins = np.histogram(sizes, bins=bins)
print(hist)
print(bins)

输出:

[3 2 3 2 2]
[ 1  3  7 10 20 22]

3. 自动计算最小等距分箱数

如果偏好等距分箱,可以先统计非零计数的唯一值数量,以此确定最小分箱数,避免空分箱:

import numpy as np

sizes = np.array([
    1, 1, 2, 3, 4, 1, 2, 4,
    9, 9, 7, 9, 10, 10, 20, 21,
])
# 统计每个值的出现次数,计算非零计数的数量
value_counts = np.bincount(sizes)
min_bins = np.sum(value_counts > 0)
# 使用最小分箱数做等距分箱
hist, bins = np.histogram(sizes, bins=min_bins)
print(hist)
print(bins)

输出:

[5 2 1 3 2]
[ 1.   5.2  9.4 13.6 17.8 22. ]

补充说明

numpy.histogram_bin_edges也支持bins='quantile'等参数,可先生成符合要求的分箱边界,再传入numpy.histogram使用,效果一致。


内容的提问来源于stack exchange,提问作者Abel Gutiérrez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 19:31:08