You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

双阈值二元分类:如何提升不确定区间最小化的优化稳定性?

问题描述

在二元分类任务中,需在真阳性率(TPR)≥p、真阴性率(TNR)≥q的约束下,最小化不确定预测的数量(即模型输出概率落在阈值a和b之间的样本数,满足0 < a < b < 1)。

当前基于Scipy的COBYLA算法实现存在以下问题:

  • 仅当初始猜测为(0.2,0.8)时能收敛到最优解,其他初始值(包括接近最优解的取值)均因无法满足约束失败;
  • 尝试SLSQP算法时,初始值不满足约束则直接终止,满足约束时返回次优解。

需要优化方案提升稳定性,摆脱对初始猜测的依赖。

问题根源分析
  1. 非光滑的目标与约束:目标函数和TPR/TNR约束依赖混淆矩阵计算,而混淆矩阵会随a、b跨过样本概率值时发生离散跳跃,导致函数不可导。COBYLA、SLSQP这类基于梯度的算法对非光滑问题鲁棒性极差,易陷入局部最优或无法找到可行域。
  2. COBYLA的局限性:该算法基于线性近似,对初始点位置敏感,处理非光滑约束时易误判可行域。
  3. SLSQP的严格性:要求初始点必须满足所有约束,否则直接终止,适用场景受限。
解决方案

以下方法可单独或组合使用,提升优化稳定性:

方法1:改用全局优化算法(推荐)

使用Scipy的differential_evolution——基于种群的全局优化算法,无需初始点,能处理非光滑目标与约束,可全局搜索最优解。

修改后代码实现

from typing import Tuple
import numpy as np
from scipy.optimize import differential_evolution
from sklearn.metrics import confusion_matrix
import pandas as pd


def objective_function(
    thresholds: Tuple[float, float],
    probs: np.ndarray,
    targets: np.ndarray,
) -> float:
    a, b = thresholds
    included_indices = (probs < a) | (probs >= b)
    # 最小化不确定样本比例(排除率)
    return 1 - float(np.mean(included_indices))


def calculate_rates(
    targets: np.ndarray,
    probs: np.ndarray,
    thresholds: Tuple[float, float],
) -> Tuple[float, float]:
    a, b = thresholds
    idx = (probs < a) | (probs >= b)
    y_true_non_excluded = targets[idx]
    y_pred_non_excluded = (probs[idx] >= b).astype(int)
    
    # 处理空样本情况
    if len(y_true_non_excluded) == 0:
        return 0.0, 0.0
    
    # 明确标签顺序,避免混淆矩阵维度错误
    tn, fp, fn, tp = confusion_matrix(
        y_true_non_excluded,
        y_pred_non_excluded,
        labels=[0,1]
    ).ravel()
    tpr = tp / (tp + fn) if (tp + fn) > 0 else 0.0
    tnr = tn / (tn + fp) if (tn + fp) > 0 else 0.0
    return tpr, tnr


def constraint_tnr(thresholds, probs, targets, p):
    tnr = calculate_rates(targets, probs, thresholds)[1]
    return tnr - p  # 约束:tnr >= p → 返回值 >=0


def constraint_tpr(thresholds, probs, targets, q):
    tpr = calculate_rates(targets, probs, thresholds)[0]
    return tpr - q  # 约束:tpr >= q → 返回值 >=0


def optimize_thresholds_de(probs: np.ndarray, targets: np.ndarray, p: float, q: float) -> Tuple[float, float]:
    # 搜索边界:添加小epsilon避免极端值
    bounds = [(1e-6, 1-1e-6), (1e-6, 1-1e-6)]
    
    constraints = [
        {'type': 'ineq', 'fun': constraint_tnr, 'args': (probs, targets, p)},
        {'type': 'ineq', 'fun': constraint_tpr, 'args': (probs, targets, q)},
        {'type': 'ineq', 'fun': lambda thrs: thrs[1] - thrs[0] - 1e-6},  # 保证a < b
    ]
    
    # 运行差分进化优化
    result = differential_evolution(
        objective_function,
        bounds=bounds,
        args=(probs, targets),
        constraints=constraints,
        popsize=15,
        maxiter=1000,
        tol=1e-6,
        seed=42  # 固定种子保证可复现
    )
    
    if result.success:
        optimal_a, optimal_b = result.x
        optimal_r = objective_function((optimal_a, optimal_b), probs, targets)
        optimal_tpr, optimal_tnr = calculate_rates(targets, probs, (optimal_a, optimal_b))
        
        print(f'Optimal a: {optimal_a:.6f}, Optimal b: {optimal_b:.6f}')
        print(f'Exclusion rate r: {optimal_r:.6f}')
        print(f'True Negative Rate: {optimal_tnr:.6f}')
        print(f'True Positive Rate: {optimal_tpr:.6f}')
    else:
        print(result)
        raise ValueError('Optimisation was not successful.')
    
    return optimal_a, optimal_b


def main() -> None:
    df = pd.read_csv('sample_data.csv')
    x = df['x'].to_numpy()
    y = df['y'].to_numpy()
    
    # 无需初始点,直接运行
    optimize_thresholds_de(x, y, p=0.9, q=0.9)

方法2:离散候选阈值搜索

最优阈值a和b必然是模型输出概率中的某个值(或相邻值),可提取所有独特概率值作为候选,遍历满足a < b的组合,筛选符合约束的候选对后选择排除率最小的。

核心代码示例

def optimize_thresholds_discrete(probs: np.ndarray, targets: np.ndarray, p: float, q: float) -> Tuple[float, float]:
    # 获取排序后的独特概率值
    unique_probs = np.sort(np.unique(probs))
    n = len(unique_probs)
    min_exclusion_rate = float('inf')
    best_a, best_b = 0.0, 1.0
    
    # 遍历所有a < b的组合
    for i in range(n):
        a = unique_probs[i]
        for j in range(i+1, n):
            b = unique_probs[j]
            tpr, tnr = calculate_rates(targets, probs, (a, b))
            if tpr >= q and tnr >= p:
                current_exclusion = objective_function((a, b), probs, targets)
                if current_exclusion < min_exclusion_rate:
                    min_exclusion_rate = current_exclusion
                    best_a, best_b = a, b
    
    if min_exclusion_rate == float('inf'):
        raise ValueError('No valid threshold pair found satisfying constraints.')
    
    print(f'Optimal a: {best_a:.6f}, Optimal b: {best_b:.6f}')
    print(f'Exclusion rate r: {min_exclusion_rate:.6f}')
    print(f'True Negative Rate: {calculate_rates(targets, probs, (best_a, best_b))[1]:.6f}')
    print(f'True Positive Rate: {calculate_rates(targets, probs, (best_a, best_b))[0]:.6f}')
    
    return best_a, best_b

该方法完全避免非光滑问题,结果可靠;但样本量大、独特概率值多时,时间复杂度为O(n²),计算量会显著增加。

方法3:改进COBYLA参数设置

若坚持使用COBYLA,可调整参数提升鲁棒性:

  • rhobeg:设置较小初始步长(如0.05),让搜索更精细;
  • maxiter:增加最大迭代次数(如1000);
  • 约束松弛:将TPR≥p改为TPR≥p-1e-3,避免浮点精度问题导致误判。

修改后的COBYLA调用代码

result = minimize(
    objective_function,
    x0=(0.5, 0.5),  # 任意初始点
    args=(probs, targets),
    method='COBYLA',
    constraints=constraints,
    options={
        'rhobeg': 0.05,
        'maxiter': 1000,
        'disp': True
    }
)
关键注意事项
  • 混淆矩阵标签顺序:指定labels=[0,1],避免因某类标签缺失导致维度错误;
  • 空样本处理:当所有样本被排除时,特殊处理TPR/TNR计算,避免除以零;
  • 数值稳定性:给阈值边界添加小epsilon(如1e-6),避免a=0、b=1或a=b的极端情况。

内容的提问来源于stack exchange,提问作者Filip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 16:34:59