Python从零实现逻辑回归:特征选择与AttributeError报错解决
问题分析与解决方案
我来帮你一步步搞定这个问题——你遇到的AttributeError是因为对特征选择的实现方式理解错了,同时我也会教你如何对齐scikit-learn的特征选择结果,让你的手动实现LR用上正确的特征。
1. 先解决AttributeError: 'list' object has no attribute 'shape'报错
你现在把X_train和X_test写成了特征名称的字符串列表,但你的Logistic Regression代码需要的是numpy数组格式的特征矩阵(每行是一个样本,每列是一个特征)。列表没有shape属性,所以执行initial_theta = np.random.rand(X_train.shape[1],1)时必然报错。
解决这个问题的核心是:从原始数据集中提取对应特征名称的列,转换成numpy数组,而不是直接写字符串列表。
2. 实现与scikit-learn一致的特征选择
要和sklearn的特征选择结果对齐,你需要先借助sklearn的特征选择工具选出特征,再把这些特征应用到你的手动LR代码中。这里以常用的SelectKBest(基于统计检验选top K特征)为例,你也可以换成RFE(递归特征消除)等其他方法:
步骤1:加载并预处理原始数据集
先把乳腺癌数据集加载成DataFrame,方便按特征名称操作:
from sklearn.datasets import load_breast_cancer import pandas as pd import numpy as np # 加载数据集 data = load_breast_cancer() df = pd.DataFrame(data.data, columns=data.feature_names) y = data.target # 标签
步骤2:用sklearn执行特征选择
from sklearn.feature_selection import SelectKBest, f_classif from sklearn.model_selection import train_test_split # 初始化特征选择器,这里选top 12个特征(和你之前手动列的数量一致) selector = SelectKBest(f_classif, k=12) selector.fit(df, y) # 获取选中的特征名称 selected_features = df.columns[selector.get_support()] print("sklearn选中的特征:", selected_features.tolist())
步骤3:划分训练/测试集并提取特征矩阵
把选中的特征转换成numpy数组格式的训练/测试集:
# 先划分训练测试集(注意要先划分再提取特征,避免数据泄露) X_train_df, X_test_df, Y_train, Y_test = train_test_split(df, y, test_size=0.2, random_state=42) # 提取选中的特征列,转成numpy数组 X_train = X_train_df[selected_features].values X_test = X_test_df[selected_features].values
步骤4:给特征矩阵添加偏置项(关键!)
手动实现Logistic Regression时,需要给特征矩阵添加一列全1的偏置项(对应theta0,即截距项),否则你的假设函数会缺少常数项:
# 添加偏置项(第一列全1) X_train = np.hstack((np.ones((X_train.shape[0], 1)), X_train)) X_test = np.hstack((np.ones((X_test.shape[0], 1)), X_test))
3. 完整修改后的代码示例
把上述步骤和你的LR代码整合,最终代码如下:
from sklearn.datasets import load_breast_cancer import pandas as pd import numpy as np from sklearn.feature_selection import SelectKBest, f_classif from sklearn.model_selection import train_test_split # ---------------------- 1. 数据加载与特征选择 ---------------------- data = load_breast_cancer() df = pd.DataFrame(data.data, columns=data.feature_names) y = data.target # 用sklearn选择top 12特征 selector = SelectKBest(f_classif, k=12) selector.fit(df, y) selected_features = df.columns[selector.get_support()] print("选中的特征:", selected_features.tolist()) # 划分训练测试集并提取特征矩阵 X_train_df, X_test_df, Y_train, Y_test = train_test_split(df, y, test_size=0.2, random_state=42) X_train = X_train_df[selected_features].values X_test = X_test_df[selected_features].values # 添加偏置项 X_train = np.hstack((np.ones((X_train.shape[0], 1)), X_train)) X_test = np.hstack((np.ones((X_test.shape[0], 1)), X_test)) # ---------------------- 2. 手动实现Logistic Regression ---------------------- def Sigmoid(z): return 1/(1 + np.exp(-z)) def Hypothesis(theta, X): return Sigmoid(X @ theta) def Cost_Function(X,Y,theta,m): hi = Hypothesis(theta, X) _y = Y.reshape(-1, 1) J = 1/float(m) * np.sum(-_y * np.log(hi) - (1-_y) * np.log(1-hi)) return J def Cost_Function_Derivative(X,Y,theta,m,alpha): hi = Hypothesis(theta,X) _y = Y.reshape(-1, 1) J = alpha/float(m) * X.T @ (hi - _y) return J def Gradient_Descent(X,Y,theta,m,alpha): new_theta = theta - Cost_Function_Derivative(X,Y,theta,m,alpha) return new_theta def Accuracy(theta): correct = 0 length = len(X_test) prediction = (Hypothesis(theta, X_test) > 0.5) _y = Y_test.reshape(-1, 1) correct = prediction == _y my_accuracy = (np.sum(correct) / length)*100 print ('LR Accuracy: ', my_accuracy, "%") def Logistic_Regression(X,Y,alpha,theta,num_iters): m = len(Y) for x in range(num_iters): new_theta = Gradient_Descent(X,Y,theta,m,alpha) theta = new_theta if x % 100 == 0: # 可选:打印中间结果 # print ('theta: ', theta) # print ('cost: ', Cost_Function(X,Y,theta,m)) pass Accuracy(theta) # 初始化参数 ep = .012 initial_theta = np.random.rand(X_train.shape[1],1) * 2 * ep - ep alpha = 0.5 iterations = 10000 # 训练并评估 Logistic_Regression(X_train,Y_train,alpha,initial_theta,iterations)
额外说明
- 如果想换成其他sklearn特征选择方法(比如RFE),只需要替换特征选择的代码块即可,核心逻辑都是先拿到选中的特征名称,再提取对应列。
- 注意一定要先划分训练测试集,再做特征选择(或者用
Pipeline),避免数据泄露影响模型泛化能力。
内容的提问来源于stack exchange,提问作者DN1
相关产品推荐
相关产品推荐

