如何通过PyCaret获取逻辑回归模型的特征估计参数?
获取PyCaret中逻辑回归的特征估计参数
PyCaret封装的模型本质是包含预处理步骤的Sklearn Pipeline,你可以通过以下步骤提取逻辑回归的特征估计参数:
1. 提取底层Sklearn逻辑回归模型
经过tune_model和optimize_threshold后,模型对象是Pipeline结构,需通过named_steps['estimator']获取最终训练好的逻辑回归实例:
sklearn_lr = optimized_lr.named_steps['estimator']
2. 获取预处理后的特征名称
由于你在setup中做了特征选择、去多重共线性等操作,需要拿到最终用于训练的特征列表,可通过get_config方法获取:
feature_names = get_config('X_train').columns.tolist()
3. 提取特征系数与截距
直接从底层模型中获取系数和截距项:
# 特征系数(逻辑回归为二分类时,coef_是二维数组,取第一个元素即可) coefficients = sklearn_lr.coef_[0] # 截距项 intercept = sklearn_lr.intercept_[0]
4. 组合成可读结果(可选)
将特征名与系数配对并排序,方便分析影响程度:
# 组合特征与系数 feature_coef = dict(zip(feature_names, coefficients)) # 按系数绝对值倒序排列 sorted_feature_coef = sorted(feature_coef.items(), key=lambda x: abs(x[1]), reverse=True) # 打印结果 print("特征估计参数(按影响程度排序):") for feature, coef in sorted_feature_coef: print(f"{feature}: {coef:.4f}") print(f"\n截距项: {intercept:.4f}")
完整代码示例
from pycaret.classification import * clf1 = setup(data = train, target = 'target', feature_selection = True, test_data = test, remove_multicollinearity = True, multicollinearity_threshold = 0.4) # create model lr = create_model('lr') # tune model tuned_lr = tune_model(lr) # optimize threshold optimized_lr = optimize_threshold(tuned_lr) # 提取底层模型 sklearn_lr = optimized_lr.named_steps['estimator'] # 获取特征名 feature_names = get_config('X_train').columns.tolist() # 获取系数和截距 coefficients = sklearn_lr.coef_[0] intercept = sklearn_lr.intercept_[0] # 组合并排序输出 feature_coef = dict(zip(feature_names, coefficients)) sorted_feature_coef = sorted(feature_coef.items(), key=lambda x: abs(x[1]), reverse=True) print("特征估计参数(按影响程度排序):") for feature, coef in sorted_feature_coef: print(f"{feature}: {coef:.4f}") print(f"\n截距项: {intercept:.4f}")
关键说明
- PyCaret的模型对象在调优后是Pipeline,必须通过
named_steps['estimator']穿透到实际训练的逻辑回归模型。 get_config('X_train')返回的是经过所有预处理后的训练特征矩阵,其列名与模型系数严格对应。
内容的提问来源于stack exchange,提问作者Gustavomoty
相关产品推荐
相关产品推荐

