You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从零实现逻辑回归:特征选择与AttributeError报错解决

问题分析与解决方案

我来帮你一步步搞定这个问题——你遇到的AttributeError是因为对特征选择的实现方式理解错了,同时我也会教你如何对齐scikit-learn的特征选择结果,让你的手动实现LR用上正确的特征。

1. 先解决AttributeError: 'list' object has no attribute 'shape'报错

你现在把X_train和X_test写成了特征名称的字符串列表,但你的Logistic Regression代码需要的是numpy数组格式的特征矩阵(每行是一个样本,每列是一个特征)。列表没有shape属性,所以执行initial_theta = np.random.rand(X_train.shape[1],1)时必然报错。

解决这个问题的核心是:从原始数据集中提取对应特征名称的列,转换成numpy数组,而不是直接写字符串列表。

2. 实现与scikit-learn一致的特征选择

要和sklearn的特征选择结果对齐,你需要先借助sklearn的特征选择工具选出特征,再把这些特征应用到你的手动LR代码中。这里以常用的SelectKBest(基于统计检验选top K特征)为例,你也可以换成RFE(递归特征消除)等其他方法:

步骤1:加载并预处理原始数据集

先把乳腺癌数据集加载成DataFrame,方便按特征名称操作:

from sklearn.datasets import load_breast_cancer
import pandas as pd
import numpy as np

# 加载数据集
data = load_breast_cancer()
df = pd.DataFrame(data.data, columns=data.feature_names)
y = data.target  # 标签

步骤2:用sklearn执行特征选择

from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.model_selection import train_test_split

# 初始化特征选择器,这里选top 12个特征(和你之前手动列的数量一致)
selector = SelectKBest(f_classif, k=12)
selector.fit(df, y)

# 获取选中的特征名称
selected_features = df.columns[selector.get_support()]
print("sklearn选中的特征:", selected_features.tolist())

步骤3:划分训练/测试集并提取特征矩阵

把选中的特征转换成numpy数组格式的训练/测试集:

# 先划分训练测试集(注意要先划分再提取特征,避免数据泄露)
X_train_df, X_test_df, Y_train, Y_test = train_test_split(df, y, test_size=0.2, random_state=42)

# 提取选中的特征列,转成numpy数组
X_train = X_train_df[selected_features].values
X_test = X_test_df[selected_features].values

步骤4:给特征矩阵添加偏置项(关键!)

手动实现Logistic Regression时,需要给特征矩阵添加一列全1的偏置项(对应theta0,即截距项),否则你的假设函数会缺少常数项:

# 添加偏置项(第一列全1)
X_train = np.hstack((np.ones((X_train.shape[0], 1)), X_train))
X_test = np.hstack((np.ones((X_test.shape[0], 1)), X_test))

3. 完整修改后的代码示例

把上述步骤和你的LR代码整合,最终代码如下:

from sklearn.datasets import load_breast_cancer
import pandas as pd
import numpy as np
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.model_selection import train_test_split

# ---------------------- 1. 数据加载与特征选择 ----------------------
data = load_breast_cancer()
df = pd.DataFrame(data.data, columns=data.feature_names)
y = data.target

# 用sklearn选择top 12特征
selector = SelectKBest(f_classif, k=12)
selector.fit(df, y)
selected_features = df.columns[selector.get_support()]
print("选中的特征:", selected_features.tolist())

# 划分训练测试集并提取特征矩阵
X_train_df, X_test_df, Y_train, Y_test = train_test_split(df, y, test_size=0.2, random_state=42)
X_train = X_train_df[selected_features].values
X_test = X_test_df[selected_features].values

# 添加偏置项
X_train = np.hstack((np.ones((X_train.shape[0], 1)), X_train))
X_test = np.hstack((np.ones((X_test.shape[0], 1)), X_test))

# ---------------------- 2. 手动实现Logistic Regression ----------------------
def Sigmoid(z):
    return 1/(1 + np.exp(-z))

def Hypothesis(theta, X):
    return Sigmoid(X @ theta)

def Cost_Function(X,Y,theta,m):
    hi = Hypothesis(theta, X)
    _y = Y.reshape(-1, 1)
    J = 1/float(m) * np.sum(-_y * np.log(hi) - (1-_y) * np.log(1-hi))
    return J

def Cost_Function_Derivative(X,Y,theta,m,alpha):
    hi = Hypothesis(theta,X)
    _y = Y.reshape(-1, 1)
    J = alpha/float(m) * X.T @ (hi - _y)
    return J

def Gradient_Descent(X,Y,theta,m,alpha):
    new_theta = theta - Cost_Function_Derivative(X,Y,theta,m,alpha)
    return new_theta

def Accuracy(theta):
    correct = 0
    length = len(X_test)
    prediction = (Hypothesis(theta, X_test) > 0.5)
    _y = Y_test.reshape(-1, 1)
    correct = prediction == _y
    my_accuracy = (np.sum(correct) / length)*100
    print ('LR Accuracy: ', my_accuracy, "%")

def Logistic_Regression(X,Y,alpha,theta,num_iters):
    m = len(Y)
    for x in range(num_iters):
        new_theta = Gradient_Descent(X,Y,theta,m,alpha)
        theta = new_theta
        if x % 100 == 0:
            # 可选:打印中间结果
            # print ('theta: ', theta)
            # print ('cost: ', Cost_Function(X,Y,theta,m))
            pass
    Accuracy(theta)

# 初始化参数
ep = .012
initial_theta = np.random.rand(X_train.shape[1],1) * 2 * ep - ep
alpha = 0.5
iterations = 10000

# 训练并评估
Logistic_Regression(X_train,Y_train,alpha,initial_theta,iterations)

额外说明

  • 如果想换成其他sklearn特征选择方法(比如RFE),只需要替换特征选择的代码块即可,核心逻辑都是先拿到选中的特征名称,再提取对应列。
  • 注意一定要先划分训练测试集,再做特征选择(或者用Pipeline),避免数据泄露影响模型泛化能力。

内容的提问来源于stack exchange,提问作者DN1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:09:56