You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python和NumPy中量化正态分布浮点数以降低误差?

高斯分布数组的最优量化方案

假设数组A的元素采样自高斯分布,需将A中每个值替换为集合R中的n_R个“代表值”之一,以最小化总量化误差。

线性量化实现(NumPy代码)

线性量化是一种可行方案,但量化误差较大,代码如下:

import numpy as np

n_A, n_R = 1_000_000, 256
mu, sig = 500, 250
A = np.random.normal(mu, sig, size = n_A)
lo, hi = np.min(A), np.max(A)
R = np.linspace(lo, hi, n_R)
I = np.round((A - lo) * (n_R - 1) / (hi - lo)).astype(np.uint32)

L = np.mean(np.abs(A - R[I]))
print('Linear loss:', L)
# 输出示例:Linspace loss: 2.3303939600700603

是否存在更优方案?可以考虑利用A的正态分布特性,或是通过迭代过程最小化“损失”函数。

基于高斯分布的加权量化方案(优化后)

调研后适配了一种加权量化方法,能得到更优的量化结果:

import numpy as np
from scipy.stats import norm

dist = norm(loc = mu, scale = sig)
bounds = dist.cdf([mu - 3*sig, mu + 3*sig])
pp = np.linspace(*bounds, n_R)
R = dist.ppf(pp)

# 找到最接近的代表值
lhits = np.clip(np.searchsorted(R, A, 'left'), 0, n_R - 1)
rhits = np.clip(np.searchsorted(R, A, 'right') - 1, 0, n_R - 1)

ldiff = R[lhits] - A
rdiff = A - R[rhits]
I = lhits
idx = np.where(rdiff < ldiff)[0]
I[idx] = rhits[idx]

L = np.mean(np.abs(A - R[I]))
print('Gaussian loss:', L)
# 输出示例:Gaussian loss: 1.6521974945326285

注:K-means聚类理论上可能得到更优结果,但针对大型数组时速度过慢,缺乏实用性。

内容的提问来源于stack exchange,提问作者Changed My Name Again

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 08:36:50