You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将独热编码离散分布平滑为峰值不变的归一化正态分布

将独热编码区间分布平滑为峰值不变的正态分布

实现思路

核心是基于独热编码的峰值位置生成正态分布概率密度,再归一化确保总和为1,通过调整标准差控制概率分散程度:

  • 定位独热编码中的峰值索引
  • 基于区间数量生成横坐标,以峰值索引为均值生成正态分布密度
  • 归一化密度值,保证总和为1

代码实现(Python + NumPy)

import numpy as np

def smooth_onehot_to_normal(onehot_arr, std_scale=0.15):
    # 定位峰值位置
    peak_idx = np.argmax(onehot_arr)
    n_bins = len(onehot_arr)
    # 将区间标准化到0-1范围,方便控制标准差
    x = np.linspace(0, 1, n_bins)
    mean = x[peak_idx]
    # std_scale控制分散程度:值越大,概率越分散到其他区间
    std = std_scale
    # 计算正态分布概率密度
    pdf = np.exp(-((x - mean)**2) / (2 * std**2))
    # 归一化,确保总和为1
    normalized_dist = pdf / pdf.sum()
    return normalized_dist

# 测试示例
test_onehot = np.array([0, 0, 1, 0, 0, 0, 0, 0, 0, 0])
smoothed_result = smooth_onehot_to_normal(test_onehot)
print("平滑后的分布:")
print(smoothed_result)
print("分布总和:", smoothed_result.sum())

参数调整说明

  • std_scale:控制概率分散程度。比如将其设为0.2时,分布会更分散;设为0.1时,分布更接近原独热编码。
  • 如果希望直接基于区间索引计算(不做0-1标准化),可以用以下变体:
def smooth_onehot_to_normal_by_idx(onehot_arr, std_factor=0.1):
    peak_idx = np.argmax(onehot_arr)
    n_bins = len(onehot_arr)
    x = np.arange(n_bins)
    mean = peak_idx
    std = n_bins * std_factor  # 基于区间总数动态设置标准差
    pdf = np.exp(-((x - mean)**2) / (2 * std**2))
    return pdf / pdf.sum()

效果验证

运行上述代码后,输出的分布会以原独热编码的1所在位置为峰值,其余区间按正态分布规律分配概率,且总和严格为1,完全符合需求。

内容的提问来源于stack exchange,提问作者Dan Jackson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 13:13:17