双阈值二元分类:如何提升不确定区间最小化的优化稳定性?
问题描述
在二元分类任务中,需在真阳性率(TPR)≥p、真阴性率(TNR)≥q的约束下,最小化不确定预测的数量(即模型输出概率落在阈值a和b之间的样本数,满足0 < a < b < 1)。
当前基于Scipy的COBYLA算法实现存在以下问题:
- 仅当初始猜测为
(0.2,0.8)时能收敛到最优解,其他初始值(包括接近最优解的取值)均因无法满足约束失败; - 尝试SLSQP算法时,初始值不满足约束则直接终止,满足约束时返回次优解。
需要优化方案提升稳定性,摆脱对初始猜测的依赖。
问题根源分析
- 非光滑的目标与约束:目标函数和TPR/TNR约束依赖混淆矩阵计算,而混淆矩阵会随
a、b跨过样本概率值时发生离散跳跃,导致函数不可导。COBYLA、SLSQP这类基于梯度的算法对非光滑问题鲁棒性极差,易陷入局部最优或无法找到可行域。 - COBYLA的局限性:该算法基于线性近似,对初始点位置敏感,处理非光滑约束时易误判可行域。
- SLSQP的严格性:要求初始点必须满足所有约束,否则直接终止,适用场景受限。
解决方案
以下方法可单独或组合使用,提升优化稳定性:
方法1:改用全局优化算法(推荐)
使用Scipy的differential_evolution——基于种群的全局优化算法,无需初始点,能处理非光滑目标与约束,可全局搜索最优解。
修改后代码实现
from typing import Tuple import numpy as np from scipy.optimize import differential_evolution from sklearn.metrics import confusion_matrix import pandas as pd def objective_function( thresholds: Tuple[float, float], probs: np.ndarray, targets: np.ndarray, ) -> float: a, b = thresholds included_indices = (probs < a) | (probs >= b) # 最小化不确定样本比例(排除率) return 1 - float(np.mean(included_indices)) def calculate_rates( targets: np.ndarray, probs: np.ndarray, thresholds: Tuple[float, float], ) -> Tuple[float, float]: a, b = thresholds idx = (probs < a) | (probs >= b) y_true_non_excluded = targets[idx] y_pred_non_excluded = (probs[idx] >= b).astype(int) # 处理空样本情况 if len(y_true_non_excluded) == 0: return 0.0, 0.0 # 明确标签顺序,避免混淆矩阵维度错误 tn, fp, fn, tp = confusion_matrix( y_true_non_excluded, y_pred_non_excluded, labels=[0,1] ).ravel() tpr = tp / (tp + fn) if (tp + fn) > 0 else 0.0 tnr = tn / (tn + fp) if (tn + fp) > 0 else 0.0 return tpr, tnr def constraint_tnr(thresholds, probs, targets, p): tnr = calculate_rates(targets, probs, thresholds)[1] return tnr - p # 约束:tnr >= p → 返回值 >=0 def constraint_tpr(thresholds, probs, targets, q): tpr = calculate_rates(targets, probs, thresholds)[0] return tpr - q # 约束:tpr >= q → 返回值 >=0 def optimize_thresholds_de(probs: np.ndarray, targets: np.ndarray, p: float, q: float) -> Tuple[float, float]: # 搜索边界:添加小epsilon避免极端值 bounds = [(1e-6, 1-1e-6), (1e-6, 1-1e-6)] constraints = [ {'type': 'ineq', 'fun': constraint_tnr, 'args': (probs, targets, p)}, {'type': 'ineq', 'fun': constraint_tpr, 'args': (probs, targets, q)}, {'type': 'ineq', 'fun': lambda thrs: thrs[1] - thrs[0] - 1e-6}, # 保证a < b ] # 运行差分进化优化 result = differential_evolution( objective_function, bounds=bounds, args=(probs, targets), constraints=constraints, popsize=15, maxiter=1000, tol=1e-6, seed=42 # 固定种子保证可复现 ) if result.success: optimal_a, optimal_b = result.x optimal_r = objective_function((optimal_a, optimal_b), probs, targets) optimal_tpr, optimal_tnr = calculate_rates(targets, probs, (optimal_a, optimal_b)) print(f'Optimal a: {optimal_a:.6f}, Optimal b: {optimal_b:.6f}') print(f'Exclusion rate r: {optimal_r:.6f}') print(f'True Negative Rate: {optimal_tnr:.6f}') print(f'True Positive Rate: {optimal_tpr:.6f}') else: print(result) raise ValueError('Optimisation was not successful.') return optimal_a, optimal_b def main() -> None: df = pd.read_csv('sample_data.csv') x = df['x'].to_numpy() y = df['y'].to_numpy() # 无需初始点,直接运行 optimize_thresholds_de(x, y, p=0.9, q=0.9)
方法2:离散候选阈值搜索
最优阈值a和b必然是模型输出概率中的某个值(或相邻值),可提取所有独特概率值作为候选,遍历满足a < b的组合,筛选符合约束的候选对后选择排除率最小的。
核心代码示例
def optimize_thresholds_discrete(probs: np.ndarray, targets: np.ndarray, p: float, q: float) -> Tuple[float, float]: # 获取排序后的独特概率值 unique_probs = np.sort(np.unique(probs)) n = len(unique_probs) min_exclusion_rate = float('inf') best_a, best_b = 0.0, 1.0 # 遍历所有a < b的组合 for i in range(n): a = unique_probs[i] for j in range(i+1, n): b = unique_probs[j] tpr, tnr = calculate_rates(targets, probs, (a, b)) if tpr >= q and tnr >= p: current_exclusion = objective_function((a, b), probs, targets) if current_exclusion < min_exclusion_rate: min_exclusion_rate = current_exclusion best_a, best_b = a, b if min_exclusion_rate == float('inf'): raise ValueError('No valid threshold pair found satisfying constraints.') print(f'Optimal a: {best_a:.6f}, Optimal b: {best_b:.6f}') print(f'Exclusion rate r: {min_exclusion_rate:.6f}') print(f'True Negative Rate: {calculate_rates(targets, probs, (best_a, best_b))[1]:.6f}') print(f'True Positive Rate: {calculate_rates(targets, probs, (best_a, best_b))[0]:.6f}') return best_a, best_b
该方法完全避免非光滑问题,结果可靠;但样本量大、独特概率值多时,时间复杂度为O(n²),计算量会显著增加。
方法3:改进COBYLA参数设置
若坚持使用COBYLA,可调整参数提升鲁棒性:
rhobeg:设置较小初始步长(如0.05),让搜索更精细;maxiter:增加最大迭代次数(如1000);- 约束松弛:将TPR≥p改为TPR≥p-1e-3,避免浮点精度问题导致误判。
修改后的COBYLA调用代码
result = minimize( objective_function, x0=(0.5, 0.5), # 任意初始点 args=(probs, targets), method='COBYLA', constraints=constraints, options={ 'rhobeg': 0.05, 'maxiter': 1000, 'disp': True } )
关键注意事项
- 混淆矩阵标签顺序:指定
labels=[0,1],避免因某类标签缺失导致维度错误; - 空样本处理:当所有样本被排除时,特殊处理TPR/TNR计算,避免除以零;
- 数值稳定性:给阈值边界添加小epsilon(如1e-6),避免
a=0、b=1或a=b的极端情况。
内容的提问来源于stack exchange,提问作者Filip
相关产品推荐
相关产品推荐

