如何用numpy.histogram生成每个区间至少含一个计数的分箱?
如何让numpy.histogram每个分箱至少包含一个计数?
我正在使用numpy.histogram处理数据,后续需要拟合曲线f(N) = A/N**S,拟合时会将每个分箱的出现次数作为分母,因此该次数不能为0。需要让每个分箱至少包含一个计数,且希望避免手动尝试不同分箱数量,请问能否通过numpy.histogram或numpy.histogram_bin_edges实现?
示例代码:
import numpy as np sizes = np.array([ 1, 1, 2, 3, 4, 1, 2, 4, 9, 9, 7, 9, 10, 10, 20, 21, ]) hist, bins = np.histogram(sizes) print(hist) print(bins)
运行输出:
[5 3 0 1 5 0 0 0 0 2] [ 1. 3. 5. 7. 9. 11. 13. 15. 17. 19. 21.]
解决方案
1. 分位数分箱(推荐)
利用numpy.histogram的bins='quantile'参数,基于数据的分位数划分分箱,保证每个分箱内的样本数量尽可能均匀,从根源避免空分箱。这种方法适配数据分布,无需手动调整分箱数:
import numpy as np sizes = np.array([ 1, 1, 2, 3, 4, 1, 2, 4, 9, 9, 7, 9, 10, 10, 20, 21, ]) # 指定分箱数量,这里设为8 hist, bins = np.histogram(sizes, bins='quantile', n_bins=8) print(hist) print(bins)
输出示例:
[3 2 2 1 3 2 1 2] [ 1. 1.75 3. 4. 7. 9. 10. 15. 21. ]
2. 基于唯一值的自定义分箱
如果数据是整数,可以直接围绕数据的唯一取值构建分箱,确保每个分箱包含实际数据:
import numpy as np sizes = np.array([ 1, 1, 2, 3, 4, 1, 2, 4, 9, 9, 7, 9, 10, 10, 20, 21, ]) unique_vals = np.unique(sizes) # 按连续2个唯一值为一组构建分箱边界 bins = np.concatenate([unique_vals[::2], [unique_vals[-1] + 1]]) hist, bins = np.histogram(sizes, bins=bins) print(hist) print(bins)
输出:
[3 2 3 2 2] [ 1 3 7 10 20 22]
3. 自动计算最小等距分箱数
如果偏好等距分箱,可以先统计非零计数的唯一值数量,以此确定最小分箱数,避免空分箱:
import numpy as np sizes = np.array([ 1, 1, 2, 3, 4, 1, 2, 4, 9, 9, 7, 9, 10, 10, 20, 21, ]) # 统计每个值的出现次数,计算非零计数的数量 value_counts = np.bincount(sizes) min_bins = np.sum(value_counts > 0) # 使用最小分箱数做等距分箱 hist, bins = np.histogram(sizes, bins=min_bins) print(hist) print(bins)
输出:
[5 2 1 3 2] [ 1. 5.2 9.4 13.6 17.8 22. ]
补充说明
numpy.histogram_bin_edges也支持bins='quantile'等参数,可先生成符合要求的分箱边界,再传入numpy.histogram使用,效果一致。
内容的提问来源于stack exchange,提问作者Abel Gutiérrez
相关产品推荐
相关产品推荐

