如何在Logistic回归特征重要性输出中显示数据集特征名称?
替换Logistic回归特征重要性输出中的序号为特征名称的方法
核心思路
要把特征序号换成实际名称,关键是拿到数据集的特征名称列表,再和系数一一对应输出即可。
具体实现
假设你用Pandas DataFrame存储训练数据(这是最常见的场景),直接从X_train的列名中提取特征名称即可,修改后的代码如下:
# 获取特征名称列表 feature_names = X_train.columns.tolist() # 获取对应类别的特征系数(此处coef_[0]对应多分类中的第一个类别) LR_importance = LR.coef_[0] # 遍历特征名与系数并输出 for name, score in zip(feature_names, LR_importance): print(f'Feature: {name}, Score: {score:.5f}')
特殊情况处理
如果你的X_train是Numpy数组(没有自带列名),需要手动定义特征名称列表,再执行同样的遍历逻辑:
# 手动定义特征名称,按数据集特征顺序填写 feature_names = ['sepal_length', 'sepal_width', 'petal_length', 'petal_width'] LR_importance = LR.coef_[0] for name, score in zip(feature_names, LR_importance): print(f'Feature: {name}, Score: {score:.5f}')
补充说明
因为你用了multi_class='multinomial',模型的coef_属性形状为(类别数量, 特征数量),coef_[0]代表第一个类别的特征系数。如果需要查看所有类别的特征重要性,可以嵌套遍历所有类别:
feature_names = X_train.columns.tolist() # 遍历每个类别 for class_idx, class_coef in enumerate(LR.coef_): print(f"\nClass {class_idx} Feature Importance:") for name, score in zip(feature_names, class_coef): print(f'Feature: {name}, Score: {score:.5f}')
内容的提问来源于stack exchange,提问作者Encipher
相关产品推荐
相关产品推荐

