You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为机器学习回归预测构建严格正的概率边界?

解决方案:非负分布替换正态分布建模球队得分概率

问题背景

当前用机器学习回归输出的球队得分预测值作为正态分布均值,历史得分标准差作为分布标准差,通过CDF计算得分区间概率,但正态分布会产生不合理的负分概率(示例中负分概率达4.7%)。需要替换为严格非负的分布,同时保留「ML预测作为核心参数+历史统计量辅助刻画分布」的优势。

原示例代码:

from scipy.stats import norm

# model predicts team will score 5 points
model_points_prediction = 5

# teams historical standard deviation of points scored is 3
team_points_standard_deviation = 3

# pass it parameters to normal distribution
dist = norm(model_points_prediction, team_points_standard_deviation)

# probability of scoring less than 0 points
print(dist.cdf(0))  # 输出约0.0478,即4.78%

可行的非负分布方案

1. 截断正态分布(推荐,改动最小)

直接对原正态分布在0处截断,将负半轴的概率全部归一化到正半轴,完全保留原正态分布的形态,且无需大幅调整参数逻辑:

from scipy.stats import truncnorm

model_points_prediction = 5
team_points_standard_deviation = 3

# 计算截断参数:a=(下界-均值)/标准差,b=(上界-均值)/标准差,这里上界设为无穷大
a = (0 - model_points_prediction) / team_points_standard_deviation
b = float('inf')

# 创建截断正态分布
dist = truncnorm(a, b, loc=model_points_prediction, scale=team_points_standard_deviation)

# 验证负分概率:直接为0
print(dist.cdf(0))  # 输出0.0

# 计算得分在0-8之间的概率
print(dist.cdf(8) - dist.cdf(0))
  • 优势:完全兼容原有的均值、标准差输入,逻辑简单,保留正态分布对得分分布的拟合特性,仅消除负分概率。

2. 对数正态分布

适用于得分的对数近似服从正态分布的场景,将ML预测的得分均值转换为对数正态分布的参数:

from scipy.stats import lognorm
import numpy as np

model_points_prediction = 5
team_points_standard_deviation = 3

# 推导对数正态分布的参数:先计算原分布的方差,再转换为对数参数
var = team_points_standard_deviation ** 2
mu = np.log(model_points_prediction ** 2 / np.sqrt(model_points_prediction ** 2 + var))
sigma = np.sqrt(np.log(1 + var / model_points_prediction ** 2))

# 创建对数正态分布(lognorm的scale参数对应exp(mu))
dist = lognorm(s=sigma, scale=np.exp(mu))

# 验证负分概率:0
print(dist.cdf(0))  # 输出0.0

# 计算得分在0-8之间的概率
print(dist.cdf(8))
  • 优势:天然非负,适合右偏的得分分布(球队得分通常符合这一特征),依然以ML预测值为核心推导分布参数。

3. 伽马分布

适合非负、右偏的连续数据,通过ML预测的均值和历史标准差推导伽马分布的形状(alpha)和尺度(beta)参数:

from scipy.stats import gamma

model_points_prediction = 5
team_points_standard_deviation = 3

# 伽马分布:均值 = alpha * beta,标准差 = sqrt(alpha) * beta
alpha = (model_points_prediction / team_points_standard_deviation) ** 2
beta = team_points_standard_deviation ** 2 / model_points_prediction

# 创建伽马分布(gamma的a参数是alpha,scale是beta)
dist = gamma(a=alpha, scale=beta)

# 验证负分概率:0
print(dist.cdf(0))  # 输出0.0

# 计算得分在0-8之间的概率
print(dist.cdf(8))
  • 优势:完全贴合非负数据特性,对右偏分布的拟合效果较好,参数推导直接基于现有ML预测值和历史统计量。

保留原方案优势的核心逻辑

以上所有方案均延续了原思路:

  • 以机器学习回归输出的得分预测值作为分布的核心位置参数(均值或由均值推导的参数)
  • 以球队历史得分的标准差作为分布的离散度参数来源
  • 通过传统统计分布的CDF计算区间概率,兼顾ML的预测能力和统计分布的概率刻画能力

内容的提问来源于stack exchange,提问作者AI92

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 23:55:17