You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在LogisticRegressionCV中获取混淆矩阵与系数矩阵的均值标准差

问题解答

1. 为什么LR.C_输出3个不同值

你使用了multi_class='ovr'(一对其余)的多分类策略,该策略会为每个类别单独训练一个二分类器(3分类任务对应3个二分类器),每个二分类器独立执行超参数C的交叉验证搜索,因此最终会返回每个类别对应的最优C值,输出长度为3的数组。

2. 如何获取最优C对应的混淆矩阵、系数均值与标准差

LogisticRegressionCV默认在全量数据拟合后仅返回全量训练的最优参数,不会自动保留每个交叉验证折的中间结果,你可以通过以下两种方案实现需求:

方案1:折内独立搜索超参数(和版本1逻辑完全对齐)

import numpy as np
from sklearn.metrics import confusion_matrix

sss = RepeatedStratifiedKFold(n_splits=K_fold, n_repeats=n_repeats ,random_state=36851234)
lambda_c = list(np.power(10.0, np.arange(-10, 3)))
scoring = 'precision_weighted'
cmn = []
coef = []

for train_index, test_index in sss.split(X,y):
    x_train, x_test = X[train_index], X[test_index]
    y_train, y_test = y[train_index], y[test_index]
    # 每折内独立执行超参数搜索
    log_reg_model = LogisticRegressionCV(max_iter=50000,fit_intercept=False,cv=K_fold,Cs=lambda_c,penalty='l1',multi_class='ovr',scoring=scoring,class_weight='balanced',solver='liblinear')
    pipe = Pipeline([('polynomial_features',polynomial), ('StandardScaler',StandardScaler()), ('logistic_regression',log_reg_model)])
    pipe.fit(x_train, y_train)
    y_pred = pipe.predict(x_test)
    # 记录当前折的系数和混淆矩阵
    LR = pipe.named_steps['logistic_regression']
    coef.append(LR.coef_)
    cmn.append(confusion_matrix(y_test,y_pred,normalize='true'))

# 计算均值与标准差
cmn_std = np.std(np.array(cmn), axis=0)
coef_std = np.std(np.array(coef), axis=0)
cmn = np.mean(np.array(cmn), axis=0)
coef = np.mean(np.array(coef), axis=0)

方案2:基于已训练模型的coefs_paths_直接提取系数统计值(无需重复训练)

如果不想重复执行超参数搜索,可以直接匹配每个类别最优C的索引,从coefs_paths_中提取所有折对应最优C的系数计算统计值:

coef_per_fold = []
# 遍历每个类别
for cls_idx in range(3):
    # 找到当前类别最优C在lambda_c列表中的索引
    best_c_idx = lambda_c.index(LR.C_[cls_idx])
    # 提取所有折、当前类别、最优C对应的系数
    cls_coef = LR.coefs_paths_[cls_idx, :, best_c_idx, :]
    coef_per_fold.append(cls_coef)

# 计算系数的均值和标准差,输出形状和版本1的结果一致为[3,6]
coef = np.mean(coef_per_fold, axis=1)
coef_std = np.std(coef_per_fold, axis=1)

混淆矩阵可直接用训练好的pipe对测试集做预测后计算,或复用方案1中每折的预测结果统计均值和标准差即可。

内容的提问来源于stack exchange,提问作者ankit agrawal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 12:27:03