You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于约束的二分类预测模型决策阈值优化问题求解

二分类模型决策阈值的成本优化问题

需求背景

我需要为已训练完成的二分类(目标标签为0和1)模型优化决策阈值——也就是模型概率预测结果被判定为阳性的临界点。举个例子:prediction is 0.7, default threshold is 0.5 -> positive prediction because 0.7 > 0.5。

我知道可以通过假阳性率这类方法确定阈值,但希望基于约束实现成本最小化来完成优化。核心目标是最小化假阳性(FP)和假阴性(FN)数量加权后的总成本,相关定义如下:

  • c_fp:假阳性的单位成本
  • c_fn:假阴性的单位成本
  • tau:决策阈值
  • f(x_i):模型对样本x_i的预测概率
  • y_i:样本x_i的真实标签
  • 还可添加假阳性/假阴性总成本上限类约束(比如threshold_fp为假阳性总成本上限)

用CVXPY建模的问题

我尝试用cvxpy构建优化问题,但核心难点是无法正确编码假阳性/假阴性的计数逻辑。初始代码如下:

import numpy as np
import cvxpy as cp

predicted_probabilities = np.array([0.3, 0.4, 0.7, 0.2, 0.8, 0.8, 0.9])
y_test = np.array([0, 0, 1, 0, 1, 0, 1])

c_fp = 100
c_fn = 1

false_positives = np.sum((predicted_probabilities >= 0.5) & (y_test == 0))
false_negatives = np.sum((predicted_probabilities < 0.5) & (y_test == 1))

current_cost = c_fp * false_positives + c_fn * false_negatives
print(f"Current cost: {current_cost}")

threshold = cp.Variable()

false_positives = cp.sum(cp.multiply((predicted_probabilities >= threshold), (y_test == 0)))

false_negatives = cp.sum(cp.multiply((predicted_probabilities < threshold), (y_test == 1)))

cost = c_fp * false_positives + c_fn * false_negatives

problem = cp.Problem(cp.Minimize(cost))

problem.solve()

print(f"Optimal threshold: {threshold.value}")

运行时会持续抛出错误:

TypeError: float() argument must be a string or a real number, not 'Inequality'

我查阅过Stack Overflow的相关问题,但场景不匹配。虽然了解整数线性规划中布尔逻辑的表达规则,但无法将其应用到当前优化场景中。

用Z3Py的解决方案

经Alex Kemper解答,最终使用z3py解决了该优化问题,代码如下:

import numpy as np
from z3 import z3

predicted_probabilities = np.array([0.3, 0.4, 0.7, 0.2, 0.8, 0.8, 0.9])
y_test = np.array([0, 0, 1, 0, 1, 0, 1])

c_fp = 100
c_fn = 10

false_positives = np.sum((predicted_probabilities >= 0.5) & (y_test == 0))
false_negatives = np.sum((predicted_probabilities < 0.5) & (y_test == 1))

current_cost = c_fp * false_positives + c_fn * false_negatives
print(f"Current cost: {current_cost}")
print(f"False positives: {false_positives}")
print(f"False negatives: {false_negatives}")

n = len(y_test)
iRange = range(n)

threshold = z3.Real("threshold")

false_positives = z3.Sum([(predicted_probabilities[i] >= threshold) for i in iRange if (y_test[i] == 0)])
false_negatives = z3.Sum([(predicted_probabilities[i] < threshold) for i in iRange if y_test[i] == 1])
cost = c_fp * false_positives + c_fn * false_negatives

opt = z3.Optimize()
opt.minimize(cost)
if opt.check() == z3.sat:
    m = opt.model()
    
    t = m[threshold].as_decimal(10)
    t = float(t)
    print(f"Optimal threshold: {t}")

    false_positives = np.sum((predicted_probabilities >= t) & (y_test == 0))
    false_negatives = np.sum((predicted_probabilities < t) & (y_test == 1)) 
    cost = c_fp * false_positives + c_fn * false_negatives
    print(f"Optimal cost: {cost}")
    print(f"False positives: {false_positives}")
    print(f"False negatives: {false_negatives}")

内容的提问来源于Stack Exchange,提问作者emil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 23:17:13