如何在Python和NumPy中量化正态分布浮点数以降低误差?
高斯分布数组的最优量化方案
假设数组A的元素采样自高斯分布,需将A中每个值替换为集合R中的n_R个“代表值”之一,以最小化总量化误差。
线性量化实现(NumPy代码)
线性量化是一种可行方案,但量化误差较大,代码如下:
import numpy as np n_A, n_R = 1_000_000, 256 mu, sig = 500, 250 A = np.random.normal(mu, sig, size = n_A) lo, hi = np.min(A), np.max(A) R = np.linspace(lo, hi, n_R) I = np.round((A - lo) * (n_R - 1) / (hi - lo)).astype(np.uint32) L = np.mean(np.abs(A - R[I])) print('Linear loss:', L) # 输出示例:Linspace loss: 2.3303939600700603
是否存在更优方案?可以考虑利用A的正态分布特性,或是通过迭代过程最小化“损失”函数。
基于高斯分布的加权量化方案(优化后)
调研后适配了一种加权量化方法,能得到更优的量化结果:
import numpy as np from scipy.stats import norm dist = norm(loc = mu, scale = sig) bounds = dist.cdf([mu - 3*sig, mu + 3*sig]) pp = np.linspace(*bounds, n_R) R = dist.ppf(pp) # 找到最接近的代表值 lhits = np.clip(np.searchsorted(R, A, 'left'), 0, n_R - 1) rhits = np.clip(np.searchsorted(R, A, 'right') - 1, 0, n_R - 1) ldiff = R[lhits] - A rdiff = A - R[rhits] I = lhits idx = np.where(rdiff < ldiff)[0] I[idx] = rhits[idx] L = np.mean(np.abs(A - R[I])) print('Gaussian loss:', L) # 输出示例:Gaussian loss: 1.6521974945326285
注:K-means聚类理论上可能得到更优结果,但针对大型数组时速度过慢,缺乏实用性。
内容的提问来源于stack exchange,提问作者Changed My Name Again
相关产品推荐
相关产品推荐

