如何为机器学习回归预测构建严格正的概率边界?
解决方案:非负分布替换正态分布建模球队得分概率
问题背景
当前用机器学习回归输出的球队得分预测值作为正态分布均值,历史得分标准差作为分布标准差,通过CDF计算得分区间概率,但正态分布会产生不合理的负分概率(示例中负分概率达4.7%)。需要替换为严格非负的分布,同时保留「ML预测作为核心参数+历史统计量辅助刻画分布」的优势。
原示例代码:
from scipy.stats import norm # model predicts team will score 5 points model_points_prediction = 5 # teams historical standard deviation of points scored is 3 team_points_standard_deviation = 3 # pass it parameters to normal distribution dist = norm(model_points_prediction, team_points_standard_deviation) # probability of scoring less than 0 points print(dist.cdf(0)) # 输出约0.0478,即4.78%
可行的非负分布方案
1. 截断正态分布(推荐,改动最小)
直接对原正态分布在0处截断,将负半轴的概率全部归一化到正半轴,完全保留原正态分布的形态,且无需大幅调整参数逻辑:
from scipy.stats import truncnorm model_points_prediction = 5 team_points_standard_deviation = 3 # 计算截断参数:a=(下界-均值)/标准差,b=(上界-均值)/标准差,这里上界设为无穷大 a = (0 - model_points_prediction) / team_points_standard_deviation b = float('inf') # 创建截断正态分布 dist = truncnorm(a, b, loc=model_points_prediction, scale=team_points_standard_deviation) # 验证负分概率:直接为0 print(dist.cdf(0)) # 输出0.0 # 计算得分在0-8之间的概率 print(dist.cdf(8) - dist.cdf(0))
- 优势:完全兼容原有的均值、标准差输入,逻辑简单,保留正态分布对得分分布的拟合特性,仅消除负分概率。
2. 对数正态分布
适用于得分的对数近似服从正态分布的场景,将ML预测的得分均值转换为对数正态分布的参数:
from scipy.stats import lognorm import numpy as np model_points_prediction = 5 team_points_standard_deviation = 3 # 推导对数正态分布的参数:先计算原分布的方差,再转换为对数参数 var = team_points_standard_deviation ** 2 mu = np.log(model_points_prediction ** 2 / np.sqrt(model_points_prediction ** 2 + var)) sigma = np.sqrt(np.log(1 + var / model_points_prediction ** 2)) # 创建对数正态分布(lognorm的scale参数对应exp(mu)) dist = lognorm(s=sigma, scale=np.exp(mu)) # 验证负分概率:0 print(dist.cdf(0)) # 输出0.0 # 计算得分在0-8之间的概率 print(dist.cdf(8))
- 优势:天然非负,适合右偏的得分分布(球队得分通常符合这一特征),依然以ML预测值为核心推导分布参数。
3. 伽马分布
适合非负、右偏的连续数据,通过ML预测的均值和历史标准差推导伽马分布的形状(alpha)和尺度(beta)参数:
from scipy.stats import gamma model_points_prediction = 5 team_points_standard_deviation = 3 # 伽马分布:均值 = alpha * beta,标准差 = sqrt(alpha) * beta alpha = (model_points_prediction / team_points_standard_deviation) ** 2 beta = team_points_standard_deviation ** 2 / model_points_prediction # 创建伽马分布(gamma的a参数是alpha,scale是beta) dist = gamma(a=alpha, scale=beta) # 验证负分概率:0 print(dist.cdf(0)) # 输出0.0 # 计算得分在0-8之间的概率 print(dist.cdf(8))
- 优势:完全贴合非负数据特性,对右偏分布的拟合效果较好,参数推导直接基于现有ML预测值和历史统计量。
保留原方案优势的核心逻辑
以上所有方案均延续了原思路:
- 以机器学习回归输出的得分预测值作为分布的核心位置参数(均值或由均值推导的参数)
- 以球队历史得分的标准差作为分布的离散度参数来源
- 通过传统统计分布的CDF计算区间概率,兼顾ML的预测能力和统计分布的概率刻画能力
内容的提问来源于stack exchange,提问作者AI92
相关产品推荐
相关产品推荐

