You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Z3Py优化二分类模型决策阈值时,优化结果准确率低于默认阈值0.5的原因排查求助

Z3Py优化二分类模型决策阈值时,优化结果准确率低于默认阈值0.5的原因排查求助

之前我在另一个问题里请教过如何优化预测模型的决策阈值,当时的方案让我开始使用z3py库。

现在我在做类似的尝试:想要优化二分类预测模型的决策阈值,以此最大化准确率。但奇怪的是,优化得到的阈值,效果居然比默认的0.5还差——按理说优化器完全可以选择0.5这个阈值,至少应该达到和默认值一样的表现才对。

我写了一个最小可复现示例(用固定随机种子生成真实标签和预测概率,确保结果可复现):

import numpy as np
from z3 import z3


def compute_eval_metrics(ground_truth, predictions):
    from sklearn.metrics import accuracy_score, f1_score

    accuracy = accuracy_score(ground_truth, predictions)
    macro_f1 = f1_score(ground_truth, predictions, average="macro")
    return accuracy, macro_f1


def optimization_acc_target(
    predictions: np.array,
    ground_truth: np.array,
    default_threshold=0.5,
):
    tp = np.sum((predictions > default_threshold) & (ground_truth == 1))
    tn = np.sum((predictions <= default_threshold) & (ground_truth == 0))

    initial_accuracy = (tp + tn) / len(ground_truth)
    print(f"Accuracy: {initial_accuracy:.3f}")

    _, initial_macro_f1_score = compute_eval_metrics(
        ground_truth, np.where(predictions > default_threshold, 1, 0)
    )

    n = len(ground_truth)
    iRange = range(n)

    threshold = z3.Real("threshold")

    opt = z3.Optimize()
    predictions = predictions.tolist()
    ground_truth = ground_truth.tolist()

    true_positives = z3.Sum(
        [
            z3.If(predictions[i] > threshold, 1, 0)
            for i in iRange
            if ground_truth[i] == 1
        ]
    )
    true_negatives = z3.Sum(
        [
            z3.If(predictions[i] <= threshold, 1, 0)
            for i in iRange
            if ground_truth[i] == 0
        ]
    )
    acc = z3.Sum(true_positives, true_negatives) / n

    # Add constraints
    opt.add(threshold >= 0.0)
    opt.add(threshold <= 1.0)

    # Maximize accuracy
    opt.maximize(acc)

    if opt.check() == z3.sat:
        m = opt.model()

        t = m[threshold].as_decimal(10)
        if type(t) == str:
            if len(t) > 1:
                t = t[:-1]
        t = float(t)
        print(f"Optimal threshold: {t}")

        optimized_accuracy, optimized_macro_f1_score = compute_eval_metrics(
            ground_truth, np.where(np.array(predictions) > t, 1, 0)
        )

        print(f"Accuracy: {optimized_accuracy:.3f} (was: {initial_accuracy:.3f})")
        print(
            f"Macro F1 Score: {optimized_macro_f1_score:.3f} (was: {initial_macro_f1_score:.3f})"
        )
        print()

    else:
        print("Failed to optimize")


np.random.seed(42)
ground_truth = np.random.randint(0, 2, size=50)
predictions = np.random.rand(50)

optimization_acc_target(
    predictions=predictions,
    ground_truth=ground_truth,
)

运行这段代码后,输出结果如下:

Accuracy: 0.600
Optimal threshold: 0.9868869366
Accuracy: 0.480 (was: 0.600)
Macro F1 Score: 0.355 (was: 0.599)

每次运行都是类似的结果:优化后的阈值效果明显比默认的0.5差。我实在搞不懂这是为什么——优化器至少应该能找到和默认阈值一样好的解吧?

为了排查问题,我试过调整z3py的写法(比如在z3.Sum里用z3.If构造),怀疑是不是数据类型不匹配导致计算错误,但调整后结果还是一样(不过这点其实也合理,因为官方示例里也是这么用的)。我还看到过一个相关的问题,但那个问题是和非线性约束有关的,我这里并没有用到非线性约束,应该不相关。

现在我很困惑,想请教大家:到底是什么原因导致优化后的阈值表现反而不如默认阈值?希望能得到排查方向、相关的背景知识或者资源指引。

备注:内容来源于stack exchange,提问作者emil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 19:20:30