You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用GridSearchCV同时优化Accuracy与F1分数?

同时优化Accuracy与F1分数的实现方法

方法1:自定义加权组合评分器

如果业务能明确两个指标的重要性占比,可以将Accuracy和F1按权重合成一个新的评分指标,以此作为网格搜索的优化目标。示例代码如下:

from sklearn.metrics import make_scorer, accuracy_score, f1_score
from sklearn.model_selection import GridSearchCV

# 自定义加权评分函数,可根据业务调整权重比例
def weighted_dual_score(y_true, y_pred):
    acc = accuracy_score(y_true, y_pred)
    f1 = f1_score(y_true, y_pred)
    # 示例:准确率和F1各占50%权重
    return 0.5 * acc + 0.5 * f1

# 转换为scikit-learn可识别的评分器
custom_scorer = make_scorer(weighted_dual_score)

# 初始化网格搜索
grid_search = GridSearchCV(
    estimator=classifier,
    param_grid=parameters,
    # 同时记录原始指标和自定义加权指标
    scoring={'accuracy': 'accuracy', 'f1': 'f1', 'weighted_dual': custom_scorer},
    refit='weighted_dual',  # 以加权指标为优化目标
    cv=5
)

grid_search.fit(X_train, y_train)

# 查看最优结果
print("最优参数组合:", grid_search.best_params_)
print("最优加权分数:", grid_search.best_score_)
print("对应准确率:", grid_search.cv_results_['mean_test_accuracy'][grid_search.best_index_])
print("对应F1分数:", grid_search.cv_results_['mean_test_f1'][grid_search.best_index_])

方法2:帕累托最优手动筛选

如果无法确定权重,可先让网格搜索计算所有参数组合的Accuracy和F1,再筛选帕累托最优组合——即不存在其他组合在两个指标上都更优的结果。示例代码:

from sklearn.model_selection import GridSearchCV

# 初始化网格搜索,关闭自动refit,仅计算所有结果
grid_search = GridSearchCV(
    estimator=classifier,
    param_grid=parameters,
    scoring=['accuracy', 'f1'],
    refit=False,
    cv=5
)

grid_search.fit(X_train, y_train)

# 提取所有参数组合的测试指标
results = grid_search.cv_results_
params_list = results['params']
mean_acc_list = results['mean_test_accuracy']
mean_f1_list = results['mean_test_f1']

# 筛选帕累托最优组合
pareto_optimal_results = []
for i in range(len(params_list)):
    is_pareto = True
    # 对比所有其他组合
    for j in range(len(params_list)):
        if i != j and mean_acc_list[j] >= mean_acc_list[i] and mean_f1_list[j] >= mean_f1_list[i]:
            is_pareto = False
            break
    if is_pareto:
        pareto_optimal_results.append({
            'params': params_list[i],
            'accuracy': round(mean_acc_list[i], 4),
            'f1': round(mean_f1_list[i], 4)
        })

# 输出帕累托最优结果
print("帕累托最优参数组合:")
for idx, res in enumerate(pareto_optimal_results):
    print(f"组合{idx+1}:")
    print(f"参数: {res['params']}")
    print(f"准确率: {res['accuracy']}, F1分数: {res['f1']}")

补充说明

scikit-learn的GridSearchCV本身不支持原生多目标优化,上述两种方法是实际业务中最常用的解决方案:

  • 加权组合适合有明确指标优先级的场景;
  • 帕累托筛选适合需要保留多个优质候选、后续人工决策的场景。

内容的提问来源于stack exchange,提问作者Carolyn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 22:15:35