You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

scikit-learn 1.0.2中LogisticRegression无feature_names_in_属性问题求助

问题描述

运行代码时出现错误:

AttributeError: 'LogisticRegression' object has no attribute 'feature_names_in_'

使用scikit-learn 1.0.2版本,尝试调用官方文档记载的feature_names_in_属性失败。完整代码及报错如下:

#imports
import numpy as np
import pandas as pd
import statistics
import scipy.sparse

from scipy.stats import chi2_contingency

from sklearn.preprocessing import FunctionTransformer, MinMaxScaler, OneHotEncoder
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.linear_model import LogisticRegression
from sklearn.impute import SimpleImputer

# train_test_split()
X_train, X_test, y_train, y_test = train_test_split(features, labels, random_state = 42)


#create functions for preprocessing

# function to replace NaN's in the ordinal and interval data 
def replace_NAN_median(X_df):
    opinions = ['opinion_seas_vacc_effective', 'opinion_seas_risk', 'opinion_seas_sick_from_vacc', 'household_adults',
                'household_children']
    for column in opinions:
        X_df[column].replace(np.nan, X_df[column].median(), inplace = True)
    return X_df

# function to replace NaN's in the catagorical data     
def replace_NAN_mode(X_df):
    miss_cat_features = ['education', 'income_poverty', 'marital_status', 'rent_or_own', 'employment_status']
    for column in miss_cat_features:
        X_df[column].replace(np.nan, statistics.mode(X_df[column]), inplace = True)
    return X_df


# Instantiate transformers
NAN_median = FunctionTransformer(replace_NAN_median)
NAN_mode = FunctionTransformer(replace_NAN_mode)

col_transformer = ColumnTransformer(transformers=
    # replace NaN's in the binary data                                
    [("NAN_0", SimpleImputer(missing_values=np.nan, strategy='constant', fill_value = 0), 
    ['behavioral_antiviral_meds', 'behavioral_avoidance','behavioral_face_mask' ,
    'behavioral_wash_hands', 'behavioral_large_gatherings', 'behavioral_outside_home',
    'behavioral_touch_face', 'doctor_recc_seasonal', 'chronic_med_condition', 
    'child_under_6_months', 'health_worker', 'health_insurance']),
    
     # MinMaxScaler on our numeric ordinal and interval data
    ("scaler", MinMaxScaler(), ['opinion_seas_vacc_effective', 'opinion_seas_risk',
                                'opinion_seas_sick_from_vacc', 
                                'household_adults', 'household_children']),
     
     # OHE catagorical string data
    ("ohe", OneHotEncoder(sparse = False), ['age_group','education', 'race', 'sex', 
                                'income_poverty', 'marital_status', 'rent_or_own',
                                'employment_status', 'census_msa'])],
     
    remainder="passthrough")


# Preprocessing Pipeline 
preprocessing_pipe = Pipeline(steps=[
    ("NAN_median", NAN_median), 
    ("NAN_mode", NAN_mode), 
    ("col_transformer", col_transformer)
    ])

# model
logreg_optimized_pipe =  Pipeline(steps=[("preprocessing_pipe", preprocessing_pipe),
                                    ("log_reg", LogisticRegression(solver = 'liblinear', random_state = 42, C = 10, penalty= 'l1'))])

#fit model to training data
logreg_optimized_pipe.fit(X_train, y_train)

#trying to get feature names
logreg_optimized_pipe.named_steps["log_reg"].feature_names_in_

报错信息:

---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
<ipython-input-38-512bfaf5962d> in <module>
----> 1 logreg_optimized_pipe.named_steps["log_reg"].feature_names_in_
       

AttributeError: 'LogisticRegression' object has no attribute 'feature_names_in_'

希望获取特征名称的可行方案。


原因分析

scikit-learn 1.0+版本的LogisticRegression确实提供feature_names_in_属性,但该属性仅当模型拟合的**输入数据包含列名信息(如pandas DataFrame)**时才会被自动赋值。你的Pipeline中,ColumnTransformer最终输出的是numpy数组(无列名),因此模型无法生成feature_names_in_属性。


解决方案

方案一:从预处理组件直接获取特征名称

这是最直接的方式,通过ColumnTransformer的get_feature_names_out()方法获取预处理后所有特征的名称,结果与模型系数一一对应:

# 获取预处理后的完整特征名称列表
feature_names = logreg_optimized_pipe.named_steps["preprocessing_pipe"].named_steps["col_transformer"].get_feature_names_out()
print(feature_names)

方案二:修改Pipeline保留列名信息

如果希望模型能直接使用feature_names_in_,可以在ColumnTransformer后添加一个转换器,将numpy数组转回带列名的DataFrame:

# 先获取预处理后的特征名
col_names = logreg_optimized_pipe.named_steps["preprocessing_pipe"].named_steps["col_transformer"].get_feature_names_out()

# 定义转换器:将数组转为DataFrame
def array_to_df(X):
    return pd.DataFrame(X, columns=col_names)

# 更新预处理Pipeline
preprocessing_pipe = Pipeline(steps=[
    ("NAN_median", NAN_median), 
    ("NAN_mode", NAN_mode), 
    ("col_transformer", col_transformer),
    ("to_df", FunctionTransformer(array_to_df))
])

# 重新定义并拟合模型
logreg_optimized_pipe = Pipeline(steps=[
    ("preprocessing_pipe", preprocessing_pipe),
    ("log_reg", LogisticRegression(solver='liblinear', random_state=42, C=10, penalty='l1'))
])
logreg_optimized_pipe.fit(X_train, y_train)

# 此时可直接调用feature_names_in_
print(logreg_optimized_pipe.named_steps["log_reg"].feature_names_in_)

内容的提问来源于stack exchange,提问作者sanderlin2013

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 19:05:23